Hardware-software design of ultra-low-power heterogeneous platforms for embedded computing

Kiamarzi, Amirhossein (2026) Hardware-software design of ultra-low-power heterogeneous platforms for embedded computing, [Dissertation thesis], Alma Mater Studiorum Università di Bologna. Dottorato di ricerca in Ingegneria e tecnologia dell'informazione per il monitoraggio strutturale e ambientale e la gestione dei rischi - eit4semm, 38 Ciclo.
Documenti full-text disponibili:
[thumbnail of Kiamarzi_Amirhossein_thesis.pdf] Documento PDF (English) - Accesso riservato fino a 31 Dicembre 2028 - Richiede un lettore di PDF come Xpdf o Adobe Acrobat Reader
Disponibile con Licenza: Creative Commons: Attribuzione 4.0 (CC BY 4.0) .
Download (6MB) | Contatta l'autore

Abstract

Embedded computing applications span from near-sensor processing to high-performance systems. Across this spectrum, there is a common trend towards heterogeneity: as a common trend, general-purpose cores coordinate specialized accelerators to meet rising performance demands within strict power budgets. This thesis investigates a reusable hardware-software stack substrate for such systems—based on open RISC-V clusters—and studies how vector engines and domain-specific units (tensor and FFT) can be co-designed with algorithms and software to improve performance per watt. The exploration starts from Parallel Ultra-Low-Power (PULP) clusters and GAP9 as representative microcontroller-class platforms and extends to systems integrating RISC-V Vector (RVV) units, coupled tensor paths, and multi-precision FFT engines under shared low-latency memory. On GAP9, we develop QR-PULP, the first open parallel QR library for embedded platforms, achieving order-of-magnitude gains in performance and energy over single-core microcontrollers. We then introduce PARSY-VDD, an end-to-end, output-only system identification framework for vibration-based structural health monitoring, validated on real infrastructures and demonstrating high accuracy with large speed and energy improvements. We curate a structural health monitoring dataset and deploy lightweight supervised and unsupervised deep models on GAP9—with and without its on-chip accelerator—to assess portability, latency, and energy trade-offs. Pushing heterogeneity further, we map a wearable ultrasound gesture-recognition pipeline to Maestro, combining RVV, tensor, and FFT units under shared memory, achieving milliwatt-class operation. Finally, we show that runtime split/merge of vector resources on a dedicated architecture (Buckbeak) enhances utilization and throughput, surpassing GAP9 in performance and energy efficiency. Together, these results provide a practical recipe—parallel RISC-V clusters as the baseline, RVV as a standard accelerator, and targeted domain engines under shared memory—along with a set of mapping rules for tiling, dataflow, and runtime reconfigurability that translates kernel efficiency into system-level performance per watt.

Abstract
Tipologia del documento
Tesi di dottorato
Autore
Kiamarzi, Amirhossein
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Multi-core Cluster, Parallel Ultra-Low-Power, Vector Accelerator
Data di discussione
27 Marzo 2026
URI

Altri metadati

Gestione del documento: Visualizza la tesi

^