Multimodal AI for human expression understanding

Mancini, Eleonora (2026) Multimodal AI for human expression understanding, [Dissertation thesis], Alma Mater Studiorum Università di Bologna. Dottorato di ricerca in Computer science and engineering, 38 Ciclo. DOI 10.48676/unibo/amsdottorato/13079.
Documenti full-text disponibili:
[thumbnail of mancini_eleonora_tesi.pdf] Documento PDF (English) - Richiede un lettore di PDF come Xpdf o Adobe Acrobat Reader
Disponibile con Licenza: Creative Commons: Attribuzione - Non Commerciale - Condividi allo Stesso Modo 4.0 (CC BY-NC-SA 4.0) .
Download (4MB)

Abstract

Human expression is inherently multimodal, spanning language, sound, movement, and other perceptual forms that together convey meaning. Each modality carries distinct yet complementary information. However, most Artificial Intelligence (AI) systems process these signals in isolation, failing to capture the complexity of human expression. This thesis investigates multimodal AI across three domains where such integration is essential: political discourse analysis, clinical assessment, and music information retrieval. It addresses three key challenges: obtaining ethically sound multimodal data, designing effective fusion architectures, and developing trustworthy explanations for domain experts. In political discourse analysis, we establish multimodal argument mining by developing the first large-scale resources combining text and speech, filling the gap left by missing audio in existing datasets. We show that prosodic and acoustic cues enhance argument detection and release an open toolkit for reproducibility. In clinical assessment, we define guidelines for responsible dataset creation and curation, and perform trimodal fusion of linguistic, acoustic, and neurophysiological signals for depression detection, achieving state-of-the-art results. Through Parkinson's detection from speech, we analyse current post-hoc explanation techniques, revealing they fail to provide clinically actionable insights and exposing a gap between algorithmic decisions and clinical understanding. In music information retrieval, we address a different data imbalance: while musical audio is abundant, lyrics are scarce or copyright-restricted, hindering multimodal approaches. We introduce WEALY, a pipeline that learns lyrics-aware representations directly from audio, enabling lyrics-based reasoning without transcriptions. Building on some of our findings, we propose a methodology for perceptually grounded interpretability across multimodal domains. We develop listenable explanations enabling experts to hear which acoustic properties drive model decisions, showing explanations are perceptually accessible and faithful to model reasoning. This thesis demonstrates that comprehensive and ethically grounded multimodal data, principled fusion strategies, and perceptually informed interpretability are fundamental for building AI systems that represent human expression in its complexity.

Abstract
Tipologia del documento
Tesi di dottorato
Autore
Mancini, Eleonora
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Multimodal AI, Natural Language Processing, Speech Processing, Multimodal Fusion, Explainable AI, Dataset Creation, AI Ethics, Argument Mining, Clinical Decision Support, Music Information Retrieval
DOI
10.48676/unibo/amsdottorato/13079
Data di discussione
26 Marzo 2026
URI

Altri metadati

Statistica sui download

Gestione del documento: Visualizza la tesi

^