Benchmarking Ultra-Low-Power μNPUs
Josh Millar, Yushan Huang, Sarab S. Sethi, Hamed Haddadi, Anil Madhavapeddy
Abstract
Efficient on-device neural network (NN) inference offers predictable latency, improved privacy and reliability, and lower operating costs for vendors than cloud-based inference. This has sparked recent development of microcontroller-scale NN accelerators, also known as neural processing units (𝜇NPUs), designed specifically for ultra-low-power applications.
We present the first comparative evaluation of a number of commercially-available 𝜇NPUs, including the first independent benchmarks for multiple platforms. To ensure fairness, we develop and open-source a model compilation pipeline supporting consistent benchmarking of quantized models across diverse microcontroller hardware. Our resulting analysis uncovers both expected performance trends as well as surprising disparities between hardware specifications and actual performance, including certain 𝜇NPUs exhibiting unexpected scaling behaviors with model complexity. This work provides a foundation for ongoing evaluation of 𝜇NPU platforms, alongside offering practical insights for both hardware and software developers in this rapidly evolving space.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd9beb0d-a878-45d8-a974-b713a4a0d09cCited by top-tier papers2
- Proxy Target: Bridging the Gap Between Discrete Spiking Neural Networks and Continuous ControlZijie Xu, Tong Bu, Zecheng Hao, Jianhao Ding et al.NeurIPS 2025 · 10 citations
- mVLM: A Vision Language Model for mNPUsZijie Chen, Guiyun Fan, Zhaoxing Yang, Rong Ding et al.CVPR 2026
Builds on8
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang et al.NeurIPS 2022 · 345 citations
- Efficient Algorithms for Device Placement of DNN Graph OperatorsJakub Tarnawski, Amar Phanishayee, Nikhil R. Devanur, Divya Mahajan et al.NeurIPS 2020 · 84 citations
- RaPiD: AI Accelerator for Ultra-low Precision Training and InferenceSwagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang, Sanchari Sen et al.ISCA 2021 · 76 citations
- MELTing Point: Mobile Evaluation of Language TransformersStefanos Laskaridis, Kleomenis Katevas, Lorenzo Minto, Hamed HaddadiMobiCom 2024 · 32 citations
Related papers
- AdaptQNet: Optimizing Quantized DNN on Microcontrollers via Adaptive Heterogeneous Processing Unit UtilizationYansong Sun, Jialuo He, Dirk Kutscher, Huangxun ChenMobiCom 2025 · 1 citation
- UDC: Unified DNAS for Compressible TinyML Models for Neural Processing UnitsIgor Fedorov, Ramon Matas Navarro, Hokchhay Tann, Chuteng Zhou et al.NeurIPS 2022 · 19 citations
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 150 citations
- Neuro-C: Neural Inference Shaped by Hardware LimitsDiletta Romano, Luca Mottola, Thiemo VoigtEuroSys 2026 · 2 citations
- QUTE: Quantifying Uncertainty in TinyML models with Early-exit-assisted ensembles for model-monitoringNikhil Pratap Ghanathe, Steven J. E. WiltonICML 2025
