Documenti full-text disponibili:
Abstract
This thesis presents a hardware–software co-design methodology for deploying deep neural networks (DNNs) on heterogeneous microcontroller (MCU) platforms for extreme-edge applications, where energy, memory, and compute budgets are tightly constrained. Object detection is used as the primary case study because it couples classification and regression, providing a representative benchmark for both discrete and continuous predictions on ultra-low-power devices. The work first analyzes how neural networks scale on resource-limited systems, quantifying accuracy–latency trade-offs and exposing the practical limits of multicore MCUs. Experiments on a nano-drone equipped with a GAP8 MCU show that model downsizing (e.g., MobileNetV2-SSD) improves latency but incurs a substantial accuracy penalty, with up to a 13% drop in mean average precision, making aggressive compression a limiting factor when multiple workloads must run on the same MCU. Motivated by these results, this thesis advocates the usage heterogeneous MCU architectures that integrate general-purpose RISC-V cores with domain-specific accelerators to increase throughput and energy efficiency. Using the GAP9 SoC as a representative heterogeneous platform, and an automated pest-detection node as a real-world exemplar, accelerator-enabled inference achieves 7.2× higher throughput and 34× lower energy than GAP8 on the same workload (MobileNetV3-SSDLite). Finally, the co-design methodology is demonstrated across two additional workloads, video object detection (VOD) and speech enhancement (SE), to validate its generality. For VOD, a multi-resolution policy (MR2-ByteTrack) interleaves low- and high-resolution frames, improving accuracy by 2.16% while reducing energy by 43.6% compared to a high-resolution-only baseline on GAP9. For SE, a mixed-precision neural beamforming pipeline combines an int8 mask estimator on the accelerator with a float32 MVDR beamformer, achieving real-time latency (20 ms) at 17.28 mJ per inference. Overall, the thesis shows that systematic co-design across algorithmic and architectural dimensions enables complex AI workloads to run efficiently on ultra-low-power heterogeneous MCUs.
Abstract
This thesis presents a hardware–software co-design methodology for deploying deep neural networks (DNNs) on heterogeneous microcontroller (MCU) platforms for extreme-edge applications, where energy, memory, and compute budgets are tightly constrained. Object detection is used as the primary case study because it couples classification and regression, providing a representative benchmark for both discrete and continuous predictions on ultra-low-power devices. The work first analyzes how neural networks scale on resource-limited systems, quantifying accuracy–latency trade-offs and exposing the practical limits of multicore MCUs. Experiments on a nano-drone equipped with a GAP8 MCU show that model downsizing (e.g., MobileNetV2-SSD) improves latency but incurs a substantial accuracy penalty, with up to a 13% drop in mean average precision, making aggressive compression a limiting factor when multiple workloads must run on the same MCU. Motivated by these results, this thesis advocates the usage heterogeneous MCU architectures that integrate general-purpose RISC-V cores with domain-specific accelerators to increase throughput and energy efficiency. Using the GAP9 SoC as a representative heterogeneous platform, and an automated pest-detection node as a real-world exemplar, accelerator-enabled inference achieves 7.2× higher throughput and 34× lower energy than GAP8 on the same workload (MobileNetV3-SSDLite). Finally, the co-design methodology is demonstrated across two additional workloads, video object detection (VOD) and speech enhancement (SE), to validate its generality. For VOD, a multi-resolution policy (MR2-ByteTrack) interleaves low- and high-resolution frames, improving accuracy by 2.16% while reducing energy by 43.6% compared to a high-resolution-only baseline on GAP9. For SE, a mixed-precision neural beamforming pipeline combines an int8 mask estimator on the accelerator with a float32 MVDR beamformer, achieving real-time latency (20 ms) at 17.28 mJ per inference. Overall, the thesis shows that systematic co-design across algorithmic and architectural dimensions enables complex AI workloads to run efficiently on ultra-low-power heterogeneous MCUs.
Tipologia del documento
Tesi di dottorato
Autore
Bompani, Luca
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Artificial Intelligence, Embedded systems, Edge AI, Heterogenous systems, Microcontrollers
DOI
10.48676/unibo/amsdottorato/12655
Data di discussione
18 Marzo 2026
URI
Altri metadati
Tipologia del documento
Tesi di dottorato
Autore
Bompani, Luca
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Artificial Intelligence, Embedded systems, Edge AI, Heterogenous systems, Microcontrollers
DOI
10.48676/unibo/amsdottorato/12655
Data di discussione
18 Marzo 2026
URI
Statistica sui download
Gestione del documento: