Miniaturized autonomous robotics, such as sub-10 cm nano-sized Unmanned Aerial Vehicles (UAVs), is a novel and pervasive technology gaining attention in both civil and industrial application domains. To enable full autonomy, tiny robots must execute multiple real-time workloads, ranging from precise Artificial Intelligence (AI)-based perception algorithms to high-throughput control pipelines, while leveraging their minimal onboard computational resources. At the same time, well-trained AI-based algorithms suffer from the domain shift problem, which leads to poor performance when deployed in real-world domains. This work presents AlSaqr, a heterogeneous and secure RISC-V System-on-Chip (SoC) that integrates a dual-core Linux-capable 64-bit host processor, an 8-core 32-bit Programmable Parallel Accelerator (PPA) embedding an FP8/FP16 tensor core, and a secure subsystem. Fabricated in GlobalFoundries 22nm technology, the proposed Pixhawk FMUv6-compliant SoC operates from 210MHz at 0.52V to 900MHz at 0.8V within a power envelope from 90mW to 635mW occupying 9mm2 silicon footprint, matching the tight size, weight, and power constraints of nano-UAVs. Post-silicon results show that the PPA achieves a peak FP16 throughput of 83GFLOp/s at 900MHz and a peak energy efficiency of 1.2 TFLOp/s/W at 210MHz. To demonstrate the SoC’s capabilities, we showcase a deep learning model for human pose estimation that produces low-level setpoints for the nano-UAV’s control loops. AlSaqr achieves a 40 frame/s inference, while concurrently fine-tuning the same model with on-device learning (i.e., backpropagation), enabling real-time on-the-fly adaptation without interruption of the primary mission.

Sinigaglia, M., Garofalo, A., Cereda, E., Tedeschi, R., Ciani, M., Tesfai, H.T., et al. (2026). A Heterogeneous System-on-Chip with an 83 GFLOp/s, 1.2 TFLOp/s/W Parallel Programmable Accelerator for Real-time On-device Learning in Autonomous Nano-UAVs. IEEE OPEN JOURNAL OF SOLID-STATE CIRCUITS, not yet assigned, 1-10 [10.1109/ojsscs.2026.3694318].

A Heterogeneous System-on-Chip with an 83 GFLOp/s, 1.2 TFLOp/s/W Parallel Programmable Accelerator for Real-time On-device Learning in Autonomous Nano-UAVs

Sinigaglia, Mattia
Primo
;
Garofalo, Angelo
Secondo
;
Tedeschi, Riccardo;Ciani, Maicol;Isachi, Victor;Palossi, Daniele;
2026

Abstract

Miniaturized autonomous robotics, such as sub-10 cm nano-sized Unmanned Aerial Vehicles (UAVs), is a novel and pervasive technology gaining attention in both civil and industrial application domains. To enable full autonomy, tiny robots must execute multiple real-time workloads, ranging from precise Artificial Intelligence (AI)-based perception algorithms to high-throughput control pipelines, while leveraging their minimal onboard computational resources. At the same time, well-trained AI-based algorithms suffer from the domain shift problem, which leads to poor performance when deployed in real-world domains. This work presents AlSaqr, a heterogeneous and secure RISC-V System-on-Chip (SoC) that integrates a dual-core Linux-capable 64-bit host processor, an 8-core 32-bit Programmable Parallel Accelerator (PPA) embedding an FP8/FP16 tensor core, and a secure subsystem. Fabricated in GlobalFoundries 22nm technology, the proposed Pixhawk FMUv6-compliant SoC operates from 210MHz at 0.52V to 900MHz at 0.8V within a power envelope from 90mW to 635mW occupying 9mm2 silicon footprint, matching the tight size, weight, and power constraints of nano-UAVs. Post-silicon results show that the PPA achieves a peak FP16 throughput of 83GFLOp/s at 900MHz and a peak energy efficiency of 1.2 TFLOp/s/W at 210MHz. To demonstrate the SoC’s capabilities, we showcase a deep learning model for human pose estimation that produces low-level setpoints for the nano-UAV’s control loops. AlSaqr achieves a 40 frame/s inference, while concurrently fine-tuning the same model with on-device learning (i.e., backpropagation), enabling real-time on-the-fly adaptation without interruption of the primary mission.
2026
Sinigaglia, M., Garofalo, A., Cereda, E., Tedeschi, R., Ciani, M., Tesfai, H.T., et al. (2026). A Heterogeneous System-on-Chip with an 83 GFLOp/s, 1.2 TFLOp/s/W Parallel Programmable Accelerator for Real-time On-device Learning in Autonomous Nano-UAVs. IEEE OPEN JOURNAL OF SOLID-STATE CIRCUITS, not yet assigned, 1-10 [10.1109/ojsscs.2026.3694318].
Sinigaglia, Mattia; Garofalo, Angelo; Cereda, Elia; Tedeschi, Riccardo; Ciani, Maicol; Tesfai, Huruy Tekle; Tolba, Mohammed; Isachi, Victor; Saleh, Ha...espandi
File in questo prodotto:
File Dimensione Formato  
A_Heterogeneous_System-on-Chip_with_an_83_GFLOp_s_1.2_TFLOp_s_W_Parallel_Programmable_Accelerator_for_Real-time_On-device_Learning_in_Autonomous_Nano-UAVs_compressed.pdf

accesso aperto

Tipo: Postprint / Author's Accepted Manuscript (AAM) - versione accettata per la pubblicazione dopo la peer-review
Licenza: Licenza per Accesso Aperto. Creative Commons Attribuzione (CCBY)
Dimensione 4.76 MB
Formato Adobe PDF
4.76 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11585/1075070
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact