Label-efficient bird's eye view perception from monocular cameras

Pineiro Monteagudo, Henrique (2026) Label-efficient bird's eye view perception from monocular cameras, [Dissertation thesis], Alma Mater Studiorum Università di Bologna. Dottorato di ricerca in Computer science and engineering, 38 Ciclo.
Documenti full-text disponibili:
[thumbnail of Henrique_PineiroMonteagudo_PhD_Thesis.pdf] Documento PDF (English) - Richiede un lettore di PDF come Xpdf o Adobe Acrobat Reader
Disponibile con Licenza: Creative Commons: Attribuzione - Non Commerciale 4.0 (CC BY-NC 4.0) .
Download (15MB)

Abstract

Assisted and automated driving systems have the potential to help prevent traffic accidents, which remain a leading cause of death worldwide. Vision-centric perception, due to its low cost, can unlock the deployment of these technologies at scale. This thesis, developed in collaboration with Verizon Connect, studies perception using bird's eye view representations extracted from monocular cameras under real-world constraints. These representations contain most of the information necessary for many tasks in driving and robotics while being compact, making them a common choice in the field. However, training machine learning models to produce them under full supervision requires expensive acquisition and annotation of three-dimensional data. To deal with a lack of annotated three-dimensional data for our use case, we propose the first method to train bird's eye view semantic segmentation networks without direct bird's eye view supervision. We introduce two variations of this method based on neural density fields and monocular depth estimation networks, and we validate it on standard benchmarks such as the nuScenes and the KITTI-360 datasets, beating baselines and showing its competitiveness as a pretraining methodology when some bird's eye view labels are available. Standard benchmark datasets used as benchmarks in bird's eye view perception are acquired under very controlled conditions and lack the diversity observed in some real-world applications, e.g., in terms of camera mounting positions. Motivated by this example, which we observe in Verizon Connect's data, we introduce a new synthetic dataset with diverse viewpoints. We study the effect of viewpoint variations in a bird's eye view semantic segmentation network trained on a single viewpoint, observing a decrease in performance. We evaluate a baseline approach combining several viewpoints at training time, observing a less pronounced decrease, suggesting that diverse training sets can help alleviate the impact of the viewpoint shift.

Abstract
Tipologia del documento
Tesi di dottorato
Autore
Pineiro Monteagudo, Henrique
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Computer vision, semantic segmentation, bird's eye view, assisted driving, 3D Computer Vision, dashcams
Data di discussione
26 Marzo 2026
URI

Altri metadati

Statistica sui download

Gestione del documento: Visualizza la tesi

^