Conference
PULP-TrainLib: Enabling On-Device Training for RISC-V Multi-core MCUs Through Performance-Driven Autotuning
| Title: | PULP-TrainLib: Enabling On-Device Training for RISC-V Multi-core MCUs Through Performance-Driven Autotuning |
|---|---|
| Authors: | Nadalini D., Rusci M., Tagliavini G., Ravaglia L., Benini L., Conti F. |
| Contributors: | Nadalini D., Rusci M., Tagliavini G., Ravaglia L., Benini L., Conti F. |
| Publisher Information: | Springer |
| Publication Year: | 2022 |
| Collection: | IRIS Università degli Studi di Bologna (CRIS - Current Research Information System) |
| Subject Terms: | artificial intelligence, computer hardware, computer network, computer programming, computer science, computer system, distributed computer system, distributed system, embedded system, engineering, internet, machine learning, microprocessor chip, parallel processing system, processor, signal processing |
| Description: | An open challenge in making Internet-of-Things sensor nodes "smart'' and self-adaptive is to enable on-chip Deep Neural Network (DNN) training on Ultra-Low-Power (ULP) microcontroller units (MCUs). To this aim, we present a framework, based on PULP-TrainLib, to deploy DNN training tasks on RISC-V-based Parallel-ULP (PULP) MCUs. PULP-TrainLib is a library of parallel software DNN primitives enabling the execution of forward and backward steps on PULP MCUs. To optimize PULP-TrainLib's kernels, we propose a strategy to automatically select and configure (autotune) the fastest among a set of tiling options and optimized floating-point matrix multiplication kernels, according to the tensor shapes of every DNN layer. Results on an 8-core RISC-V MCU show that our auto-tuned primitives improve MAC/clk by up to 2.4x compared to "one-size-fits-all'' matrix multiplication, achieving up to 4.39 MAC/clk - 36.6x better than a commercial STM32L4 MCU executing the same DNN layer training workload. Furthermore, our strategy proves to be 30.7x faster than AIfES, a state-of-the-art training library for MCUs, while training a complete TinyML model. |
| Document Type: | conference object |
| File Description: | STAMPA |
| Language: | English |
| Relation: | info:eu-repo/semantics/altIdentifier/isbn/978-3-031-15073-9; info:eu-repo/semantics/altIdentifier/isbn/978-3-031-15074-6; info:eu-repo/semantics/altIdentifier/wos/WOS:000874744300013; ispartofbook:Embedded Computer Systems: Architectures, Modeling, and Simulation; 22nd International Conference, SAMOS 2022; volume:13511; firstpage:200; lastpage:216; numberofpages:17; info:eu-repo/grantAgreement/EC/H2020/826060, 863337; https://hdl.handle.net/11585/900686 |
| DOI: | 10.1007/978-3-031-15074-6_13 |
| Availability: | https://hdl.handle.net/11585/900686 https://doi.org/10.1007/978-3-031-15074-6_13 https://link.springer.com/chapter/10.1007/978-3-031-15074-6_13 |
| Rights: | info:eu-repo/semantics/openAccess |
| Accession Number: | edsbas.BF3FAEEB |
| Database: | BASE |
| DOI: | 10.1007/978-3-031-15074-6_13 |
|---|