Academic Journal
EmbBERT: Attention under 2 MB memory.
| Τίτλος: | EmbBERT: Attention under 2 MB memory. |
|---|---|
| Συγγραφείς: | Bravin R; Department of Electronics, Information and Bioengineering, Politecnico di Milano, Via Ponzio 34/5, Milano, 20133, Italy. Electronic address: riccardo.bravin@polimi.it., Pavan M; Department of Electronics, Information and Bioengineering, Politecnico di Milano, Via Ponzio 34/5, Milano, 20133, Italy. Electronic address: massimo.pavan@polimi.it., Hesham Yousef Shalby H; Department of Electronics, Information and Bioengineering, Politecnico di Milano, Via Ponzio 34/5, Milano, 20133, Italy. Electronic address: hazemhesham.shalby@polimi.it., Pittorino F; Department of Electronics, Information and Bioengineering, Politecnico di Milano, Via Ponzio 34/5, Milano, 20133, Italy. Electronic address: fabrizio.pittorino@polimi.it., Roveri M; Department of Electronics, Information and Bioengineering, Politecnico di Milano, Via Ponzio 34/5, Milano, 20133, Italy. Electronic address: manuel.roveri@polimi.it. |
| Πηγή: | Neural networks : the official journal of the International Neural Network Society [Neural Netw] 2026 Aug; Vol. 200, pp. 108800. Date of Electronic Publication: 2026 Mar 06. |
| Τύπος έκδοσης: | Journal Article |
| Γλώσσα: | English |
| Στοιχεία περιοδικού: | Publisher: Pergamon Press Country of Publication: United States NLM ID: 8805018 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1879-2782 (Electronic) Linking ISSN: 08936080 NLM ISO Abbreviation: Neural Netw Subsets: MEDLINE |
| Imprint Name(s): | Original Publication: New York : Pergamon Press, [c1988- |
| Ιατρικοί όροι (MeSH): | Natural Language Processing* , Attention* , Neural Networks, Computer* , Memory*, Humans |
| Περίληψη: | Transformer architectures based on the attention mechanism have revolutionized natural language processing (NLP), driving major breakthroughs across virtually every NLP task. However, their substantial memory and computational requirements still hinder deployment on ultra-constrained devices such as wearables and Internet-of-Things (IoT) units, where available memory is limited to just a few megabytes. To address this challenge, we introduce EmbBERT, a tiny language model (TLM) architecturally designed for extreme efficiency. The model integrates a compact embedding layer, streamlined feed-forward blocks, and an efficient attention mechanism that together enable optimal performance under strict memory budgets. Through this redesign for the extreme edge, we demonstrate that highly simplified transformer architectures remain remarkably effective under tight resource constraints. EmbBERT requires only 2 MB of total memory, and achieves accuracy performance comparable to the ones of state-of-the-art (SotA) models that require a 10 × memory budget. Extensive experiments on the curated TinyNLP benchmark and the GLUE suite confirm that EmbBERT achieves competitive accuracy, comparable to that of larger SotA models, and consistently outperforms downsized versions of BERT and MAMBA of similar size. Furthermore, we demonstrate the model's resilience to 8-bit quantization, which further reduces memory usage to just 781 kB, and the scalability of the EmbBERT architecture across the sub-megabyte to tens-of-megabytes range. Finally, we perform an ablation study demonstrating the positive contributions of all components and the pre-training procedure. All code, scripts, and checkpoints are publicly released to ensure reproducibility: https://github.com/RiccardoBravin/tiny-LLM. (Copyright © 2026 The Authors. Published by Elsevier Ltd.. All rights reserved.) |
| Competing Interests: | Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. |
| Contributed Indexing: | Keywords: Efficient deep learning; Hardware acceleration; Language models; Model compression; Natural language processing; Tiny machine learning |
| Entry Date(s): | Date Created: 20260312 Date Completed: 20260711 Latest Revision: 20260711 |
| Update Code: | 20260711 |
| DOI: | 10.1016/j.neunet.2026.108800 |
| PMID: | 41819621 |
| Βάση Δεδομένων: | MEDLINE |
| ISSN: | 1879-2782 |
|---|---|
| DOI: | 10.1016/j.neunet.2026.108800 |