Academic Journal

A 23-µJ-per-Frame All-on-Chip TinyML U-Net Processor for Real-Time Autonomous Image Segmentation in Miniaturized Ultrasound Devices.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: A 23-µJ-per-Frame All-on-Chip TinyML U-Net Processor for Real-Time Autonomous Image Segmentation in Miniaturized Ultrasound Devices.
Συγγραφείς: Song Z, Guler U, Chandrakasan A
Πηγή: IEEE transactions on biomedical circuits and systems [IEEE Trans Biomed Circuits Syst] 2026 Aug; Vol. 20 (4), pp. 540-550.
Τύπος έκδοσης: Journal Article
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: IEEE Country of Publication: United States NLM ID: 101312520 Publication Model: Print Cited Medium: Internet ISSN: 1940-9990 (Electronic) Linking ISSN: 19324545 NLM ISO Abbreviation: IEEE Trans Biomed Circuits Syst Subsets: MEDLINE
Imprint Name(s): Original Publication: New York, NY : IEEE, c2007-
Ιατρικοί όροι (MeSH): Image Processing, Computer-Assisted*/methods , Image Processing, Computer-Assisted*/instrumentation , Neural Networks, Computer*, Ultrasonography/instrumentation ; Ultrasonography/methods ; Humans ; Algorithms
Περίληψη: Autonomous medical image segmentation enables critical applications, including urinary retention monitoring, prenatal fetal biometry, neuromodulation, and cardiovascular monitoring. Its deployment in wearable ultrasound patches demands on-device processing to preserve patient privacy and enable operation beyond clinical facilities. U-Net achieves state-of-the-art performance for biomedical segmentation, and recent binarized U-Nets retain high clinical accuracy with dramatically reduced computational cost. However, existing binary neural network (BNN) accelerators cannot support medical-grade segmentation due to missing accuracy-enhancing features, poor hardware utilization for compute-optimal layers, and memory bottlenecks requiring costly external DRAM. This work presents a 0.81 mm2 fully-integrated U-Net processor in 28nm featuring: 1) mixed-precision datapaths combining binary convolution with 4-bit skip connections for clinical accuracy; 2) systematic design space exploration across 9,390 configurations optimizing energy-latency tradeoffs; 3) interleaved memory representation and halo reuse for energy-efficient battery-powered operation; and 4) hardware-supported layer fusion and lossless compression eliminating external memory while reducing peak on-chip usage by 3.16x and 1.38x, respectively. Validated on bladder and fetal head segmentation datasets, the processor achieves 13.4 frames per second (fps) and 23 µJ per frame, enabling real-time autonomous monitoring in wearable medical devices.
Entry Date(s): Date Created: 20260323 Date Completed: 20260729 Latest Revision: 20260730
Update Code: 20260730
DOI: 10.1109/TBCAS.2026.3676752
PMID: 41870933
Βάση Δεδομένων: MEDLINE