Academic Journal
A 23-µJ-per-Frame All-on-Chip TinyML U-Net Processor for Real-Time Autonomous Image Segmentation in Miniaturized Ultrasound Devices.
| Τίτλος: | A 23-µJ-per-Frame All-on-Chip TinyML U-Net Processor for Real-Time Autonomous Image Segmentation in Miniaturized Ultrasound Devices. |
|---|---|
| Συγγραφείς: | Song Z, Guler U, Chandrakasan A |
| Πηγή: | IEEE transactions on biomedical circuits and systems [IEEE Trans Biomed Circuits Syst] 2026 Aug; Vol. 20 (4), pp. 540-550. |
| Τύπος έκδοσης: | Journal Article |
| Γλώσσα: | English |
| Στοιχεία περιοδικού: | Publisher: IEEE Country of Publication: United States NLM ID: 101312520 Publication Model: Print Cited Medium: Internet ISSN: 1940-9990 (Electronic) Linking ISSN: 19324545 NLM ISO Abbreviation: IEEE Trans Biomed Circuits Syst Subsets: MEDLINE |
| Imprint Name(s): | Original Publication: New York, NY : IEEE, c2007- |
| Ιατρικοί όροι (MeSH): | Image Processing, Computer-Assisted*/methods , Image Processing, Computer-Assisted*/instrumentation , Neural Networks, Computer*, Ultrasonography/instrumentation ; Ultrasonography/methods ; Humans ; Algorithms |
| Περίληψη: | Autonomous medical image segmentation enables critical applications, including urinary retention monitoring, prenatal fetal biometry, neuromodulation, and cardiovascular monitoring. Its deployment in wearable ultrasound patches demands on-device processing to preserve patient privacy and enable operation beyond clinical facilities. U-Net achieves state-of-the-art performance for biomedical segmentation, and recent binarized U-Nets retain high clinical accuracy with dramatically reduced computational cost. However, existing binary neural network (BNN) accelerators cannot support medical-grade segmentation due to missing accuracy-enhancing features, poor hardware utilization for compute-optimal layers, and memory bottlenecks requiring costly external DRAM. This work presents a 0.81 mm2 fully-integrated U-Net processor in 28nm featuring: 1) mixed-precision datapaths combining binary convolution with 4-bit skip connections for clinical accuracy; 2) systematic design space exploration across 9,390 configurations optimizing energy-latency tradeoffs; 3) interleaved memory representation and halo reuse for energy-efficient battery-powered operation; and 4) hardware-supported layer fusion and lossless compression eliminating external memory while reducing peak on-chip usage by 3.16x and 1.38x, respectively. Validated on bladder and fetal head segmentation datasets, the processor achieves 13.4 frames per second (fps) and 23 µJ per frame, enabling real-time autonomous monitoring in wearable medical devices. |
| Entry Date(s): | Date Created: 20260323 Date Completed: 20260729 Latest Revision: 20260730 |
| Update Code: | 20260730 |
| DOI: | 10.1109/TBCAS.2026.3676752 |
| PMID: | 41870933 |
| Βάση Δεδομένων: | MEDLINE |
καταχωρήστε σχόλιο πρώτοι!