Neural image synthesis for industrial computer vision applications

Ali, Musawar (2026) Neural image synthesis for industrial computer vision applications, [Dissertation thesis], Alma Mater Studiorum Università di Bologna. Dottorato di ricerca in Computer science and engineering, 38 Ciclo. DOI 10.48676/unibo/amsdottorato/12762.
Documenti full-text disponibili:
[thumbnail of Final_Merged_Thesis.pdf] Documento PDF (English) - Richiede un lettore di PDF come Xpdf o Adobe Acrobat Reader
Disponibile con Licenza: Creative Commons: Attribuzione 4.0 (CC BY 4.0) .
Download (171MB)

Abstract

Neural image synthesis has advanced significantly in recent years, driven by powerful generative models such as Generative Adversarial Networks (GANs), Denoising Diffusion Probabilistic Models (DDPMs), Latent Diffusion Models (LDMs), Neural Radiance Fields (NeRFs), and Gaussian Splatting. Among these, LDMs have delivered strong capabilities in generating high-fidelity, realistic images across diverse domains with applications ranging from artistic creation and data augmentation to the generation of rare events. These methods are being adopted in industrial settings, where synthetic data helps mitigate challenges associated with data scarcity. For example, to generate synthetic industrial defects for training anomaly inspection models, where collecting real defective samples is often impractical due to cost. Neural image synthesis is a promising tool to generate synthetic anomaly images for industrial quality control. Additionally, NeRFs and Gaussian Splatting have shown significant progress for novel view synthesis, where accurate camera poses are necessary for rendering realistic viewpoints. This capability aligns with industrial requirements, as many applications demand reliable methods for synthesizing handheld objects from multiple viewpoints. These advances highlight the growing role of neural image synthesis in both research and industrial practices. Following these lines of research, this thesis first demonstrates how synthetic images can improve downstream anomaly inspection through stable diffusion inpainting using AnomalyControl and VarianceGenerator. Then a benchmark for novel view synthesis of handheld objects is developed and presented. Through experiments with adapted baselines, we show the need for more accurate pose estimation techniques in scenarios involving partial hand occlusions and particularly in setups where only RGB images are available.

Abstract
Tipologia del documento
Tesi di dottorato
Autore
Ali, Musawar
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Deep Learning, Industrial Computer Vision, 3D Vision
DOI
10.48676/unibo/amsdottorato/12762
Data di discussione
26 Marzo 2026
URI

Altri metadati

Statistica sui download

Gestione del documento: Visualizza la tesi

^