Documenti full-text disponibili:
Abstract
This thesis presents a hardware–software co-design methodology for deploying deep neural networks (DNNs) on heterogeneous microcontroller (MCU) platforms for extreme-edge applications, where energy, memory, and compute budgets are tightly constrained. Object detection is used as the primary case study because it couples classification and regression, providing a representative benchmark for both discrete and continuous predictions on ultra-low-power devices. The work first analyzes how neural networks scale on resource-limited systems, quantifying accuracy–latency trade-offs and exposing the practical limits of multicore MCUs. Experiments on a nano-drone equipped with a GAP8 MCU show that model downsizing (e.g., MobileNetV2-SSD) improves latency but incurs a substantial accuracy penalty, with up to a 13% drop in mean average precision, making aggressive compression a limiting factor when multiple workloads must run on the same MCU. Motivated by these results, this thesis advocates the usage heterogeneous MCU architectures that integrate general-purpose RISC-V cores with domain-specific accelerators to increase throughput and energy efficiency. Using the GAP9 SoC as a representative heterogeneous platform, and an automated pest-detection node as a real-world exemplar, accelerator-enabled inference achieves 7.2× higher throughput and 34× lower energy than GAP8 on the same workload (MobileNetV3-SSDLite). Finally, the co-design methodology is demonstrated across two additional workloads, video object detection (VOD) and speech enhancement (SE), to validate its generality. For VOD, a multi-resolution policy (MR2-ByteTrack) interleaves low- and high-resolution frames, improving accuracy by 2.16% while reducing energy by 43.6% compared to a high-resolution-only baseline on GAP9. For SE, a mixed-precision neural beamforming pipeline combines an int8 mask estimator on the accelerator with a float32 MVDR beamformer, achieving real-time latency (20 ms) at 17.28 mJ per inference. Overall, the thesis shows that systematic co-design across algorithmic and architectural dimensions enables complex AI workloads to run efficiently on ultra-low-power heterogeneous MCUs.
Abstract
This thesis presents a hardware–software co-design methodology for deploying deep neural networks (DNNs) on heterogeneous microcontroller (MCU) platforms for extreme-edge applications, where energy, memory, and compute budgets are tightly constrained. Object detection is used as the primary case study because it couples classification and regression, providing a representative benchmark for both discrete and continuous predictions on ultra-low-power devices. The work first analyzes how neural networks scale on resource-limited systems, quantifying accuracy–latency trade-offs and exposing the practical limits of multicore MCUs. Experiments on a nano-drone equipped with a GAP8 MCU show that model downsizing (e.g., MobileNetV2-SSD) improves latency but incurs a substantial accuracy penalty, with up to a 13% drop in mean average precision, making aggressive compression a limiting factor when multiple workloads must run on the same MCU. Motivated by these results, this thesis advocates the usage heterogeneous MCU architectures that integrate general-purpose RISC-V cores with domain-specific accelerators to increase throughput and energy efficiency. Using the GAP9 SoC as a representative heterogeneous platform, and an automated pest-detection node as a real-world exemplar, accelerator-enabled inference achieves 7.2× higher throughput and 34× lower energy than GAP8 on the same workload (MobileNetV3-SSDLite). Finally, the co-design methodology is demonstrated across two additional workloads, video object detection (VOD) and speech enhancement (SE), to validate its generality. For VOD, a multi-resolution policy (MR2-ByteTrack) interleaves low- and high-resolution frames, improving accuracy by 2.16% while reducing energy by 43.6% compared to a high-resolution-only baseline on GAP9. For SE, a mixed-precision neural beamforming pipeline combines an int8 mask estimator on the accelerator with a float32 MVDR beamformer, achieving real-time latency (20 ms) at 17.28 mJ per inference. Overall, the thesis shows that systematic co-design across algorithmic and architectural dimensions enables complex AI workloads to run efficiently on ultra-low-power heterogeneous MCUs.
Tipologia del documento
Tesi di dottorato
Autore
Bompani, Luca
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Artificial Intelligence, Embedded systems, Edge AI, Heterogenous systems, Microcontrollers
Data di discussione
18 Marzo 2026
URI
Altri metadati
Tipologia del documento
Tesi di dottorato
Autore
Bompani, Luca
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Artificial Intelligence, Embedded systems, Edge AI, Heterogenous systems, Microcontrollers
Data di discussione
18 Marzo 2026
URI
Statistica sui download
Gestione del documento: