Mirsalari, Seyed Ahmad
(2026)
Leveraging transprecision computing techniques for Ultra-Low-Power IoT devices, [Dissertation thesis], Alma Mater Studiorum Università di Bologna.
Dottorato di ricerca in
Data science and computation, 37 Ciclo. DOI 10.48676/unibo/amsdottorato/12717.
Documenti full-text disponibili:
Abstract
This dissertation investigates how transprecision computing—the use of multiple numeric precisions across the hardware-software stack—enables real-time machine learning (ML) and digital signal processing (DSP) workloads on Internet of Things (IoT) edge devices based on microcontroller units (MCUs). Motivated by strict energy and memory constraints, the work targets floating-point “smallFloats” formats (FP16/BF16/FP8) as a middle ground between fixed-point and full-precision arithmetic, preserving application-level quality while substantially reducing computation and data footprint. The thesis presents four main contributions: (1) the adoption of FP8 beyond deep learning (DL), (2) TransLib, a transprecision kernel library for embedded multi-core platforms, (3) an optimized Self-Organizing Map (SOM) design for on-device bacterial-genome identification, and (4) StreamEase, an automatic streaming transformation for Temporal Convolutional Networks (TCNs). First, the thesis explores FP8 formats outside the DL domain across classic DSP/ML kernels. With proper exponent allocation and scaling strategies, FP8 achieves acceptable numerical error relative to FP32 while delivering substantial speedups of 3.14×-18.81× on 1-8 cores. Second, TransLib provides a workflow for transprecision design on embedded multi-core systems, supporting datatype exploration, fixed- and mixed-precision mappings, memory-footprint analysis, and an end-to-end heart-rate detection pipeline, enabling systematic tuning of accuracy–performance trade-off. Third, a joint algorithm–architecture co-design exploits FP8 arithmetic to enable ultra-low-power SOM-based genomics. This approach achieves accurate DNA-sequence classification with a 2× memory reduction versus 16-bit representations, speedups up to 18.7× over FP32, and a 6.72× energy-efficiency improvement over FP32 (vs. 3.15× for a CGRA baseline). Finally, StreamEase automatically converts non-streaming TCNs into streaming networks, enabling real-time inference with multi-timestep processing while preserving accuracy. On GAP9, a Conv-TasNet speech-enhancement model reaches 2 ms inference latency—33% of a 6.25 ms budget—with 108.9× fewer MACs and 27.7× fewer cycles than the non-streaming baseline. Collectively, this dissertation provides a hardware-validated methodology for deploying transprecision DSP/ML applications on MCU-class edge devices.
Abstract
This dissertation investigates how transprecision computing—the use of multiple numeric precisions across the hardware-software stack—enables real-time machine learning (ML) and digital signal processing (DSP) workloads on Internet of Things (IoT) edge devices based on microcontroller units (MCUs). Motivated by strict energy and memory constraints, the work targets floating-point “smallFloats” formats (FP16/BF16/FP8) as a middle ground between fixed-point and full-precision arithmetic, preserving application-level quality while substantially reducing computation and data footprint. The thesis presents four main contributions: (1) the adoption of FP8 beyond deep learning (DL), (2) TransLib, a transprecision kernel library for embedded multi-core platforms, (3) an optimized Self-Organizing Map (SOM) design for on-device bacterial-genome identification, and (4) StreamEase, an automatic streaming transformation for Temporal Convolutional Networks (TCNs). First, the thesis explores FP8 formats outside the DL domain across classic DSP/ML kernels. With proper exponent allocation and scaling strategies, FP8 achieves acceptable numerical error relative to FP32 while delivering substantial speedups of 3.14×-18.81× on 1-8 cores. Second, TransLib provides a workflow for transprecision design on embedded multi-core systems, supporting datatype exploration, fixed- and mixed-precision mappings, memory-footprint analysis, and an end-to-end heart-rate detection pipeline, enabling systematic tuning of accuracy–performance trade-off. Third, a joint algorithm–architecture co-design exploits FP8 arithmetic to enable ultra-low-power SOM-based genomics. This approach achieves accurate DNA-sequence classification with a 2× memory reduction versus 16-bit representations, speedups up to 18.7× over FP32, and a 6.72× energy-efficiency improvement over FP32 (vs. 3.15× for a CGRA baseline). Finally, StreamEase automatically converts non-streaming TCNs into streaming networks, enabling real-time inference with multi-timestep processing while preserving accuracy. On GAP9, a Conv-TasNet speech-enhancement model reaches 2 ms inference latency—33% of a 6.25 ms budget—with 108.9× fewer MACs and 27.7× fewer cycles than the non-streaming baseline. Collectively, this dissertation provides a hardware-validated methodology for deploying transprecision DSP/ML applications on MCU-class edge devices.
Tipologia del documento
Tesi di dottorato
Autore
Mirsalari, Seyed Ahmad
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
37
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Transprecision computing; low-precision floating-point arithmetic; embedded multi-core systems; real-time DSP and ML; edge AI; streaming neural networks; energy-efficient computing
DOI
10.48676/unibo/amsdottorato/12717
Data di discussione
25 Marzo 2026
URI
Altri metadati
Tipologia del documento
Tesi di dottorato
Autore
Mirsalari, Seyed Ahmad
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
37
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Transprecision computing; low-precision floating-point arithmetic; embedded multi-core systems; real-time DSP and ML; edge AI; streaming neural networks; energy-efficient computing
DOI
10.48676/unibo/amsdottorato/12717
Data di discussione
25 Marzo 2026
URI
Statistica sui download
Gestione del documento: