Multimodal AI for 3D-language understanding

Amaduzzi, Andrea (2026) Multimodal AI for 3D-language understanding, [Dissertation thesis], Alma Mater Studiorum Università di Bologna. Dottorato di ricerca in Data science and computation, 37 Ciclo. DOI 10.48676/unibo/amsdottorato/12771.
Documenti full-text disponibili:
[thumbnail of amaduzzi_andrea_tesi.pdf] Documento PDF (English) - Richiede un lettore di PDF come Xpdf o Adobe Acrobat Reader
Disponibile con Licenza: Creative Commons: Attribuzione 4.0 (CC BY 4.0) .
Download (31MB)

Abstract

Understanding the world through natural language is a fundamental challenge in computer vision, requiring models to bridge the gap between geometric representations and semantic understanding. While recent advances in deep learning have enabled impressive progress in multimodal tasks, effectively connecting 3D data with language remains difficult due to the computational complexity of 3D representations and the scarcity of high-quality paired 3D-language datasets. In this Thesis, we focus on developing novel deep learning techniques that can efficiently process and describe 3D data through natural language. First, we address the critical lack of high quality 3D-language datasets and reliable evaluation methods for text-to-3D generation by introducing a benchmark comprising GPT2Shape, a dataset of 3D models with high-quality language annotations, CrossCoherence, a novel quantitative metric for evaluating text-to-3D coherence, and HST, a human-validated test set. This benchmark demonstrates superior performance compared to existing 2D-based evaluation metrics. In the second part of this Thesis, we introduce NeRF-language models — a new paradigm of Multimodal Large Language Models that directly process the weights of Neural Radiance Fields (NeRFs) rather than explicit representations like images or point clouds materialized from them. We show that this approach enables holistic 3D understanding without computationally expensive rendering, achieving state-of-the-art performance on the first NeRF-language datasets. Finally, we propose a method for computing spatially-aware tokens directly from the weights of a NeRF, enabling fine-grained spatial understanding and detailed scene comprehension. Our resulting model demonstrates strong capabilities in spatial reasoning on multi-object scenes, with promising generalization to real-world scenarios.

Abstract
Tipologia del documento
Tesi di dottorato
Autore
Amaduzzi, Andrea
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
37
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
LLM, MLLM, Large Language Models, Multimodal Large Language Models, NeRF, Computer Vision, AI
DOI
10.48676/unibo/amsdottorato/12771
Data di discussione
25 Marzo 2026
URI

Altri metadati

Statistica sui download

Gestione del documento: Visualizza la tesi

^