Benchmarking Ultra-Low-Power μNPUs
Josh Millar, Yushan Huang, Sarab S. Sethi, Hamed Haddadi, Anil Madhavapeddy
摘要
Efficient on-device neural network (NN) inference offers predictable latency, improved privacy and reliability, and lower operating costs for vendors than cloud-based inference. This has sparked recent development of microcontroller-scale NN accelerators, also known as neural processing units (𝜇NPUs), designed specifically for ultra-low-power applications.
We present the first comparative evaluation of a number of commercially-available 𝜇NPUs, including the first independent benchmarks for multiple platforms. To ensure fairness, we develop and open-source a model compilation pipeline supporting consistent benchmarking of quantized models across diverse microcontroller hardware. Our resulting analysis uncovers both expected performance trends as well as surprising disparities between hardware specifications and actual performance, including certain 𝜇NPUs exhibiting unexpected scaling behaviors with model complexity. This work provides a foundation for ongoing evaluation of 𝜇NPU platforms, alongside offering practical insights for both hardware and software developers in this rapidly evolving space.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Proxy Target: Bridging the Gap Between Discrete Spiking Neural Networks and Continuous ControlZijie Xu, Tong Bu, Zecheng Hao, Jianhao Ding 等NeurIPS 2025 · 被引用 10 次
- mVLM: A Vision Language Model for mNPUsZijie Chen, Guiyun Fan, Zhaoxing Yang, Rong Ding 等CVPR 2026
它引用的顶会 Paper8
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang 等NeurIPS 2022 · 被引用 345 次
- Efficient Algorithms for Device Placement of DNN Graph OperatorsJakub Tarnawski, Amar Phanishayee, Nikhil R. Devanur, Divya Mahajan 等NeurIPS 2020 · 被引用 84 次
- RaPiD: AI Accelerator for Ultra-low Precision Training and InferenceSwagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang, Sanchari Sen 等ISCA 2021 · 被引用 76 次
- MELTing Point: Mobile Evaluation of Language TransformersStefanos Laskaridis, Kleomenis Katevas, Lorenzo Minto, Hamed HaddadiMobiCom 2024 · 被引用 32 次
相关 Paper
- AdaptQNet: Optimizing Quantized DNN on Microcontrollers via Adaptive Heterogeneous Processing Unit UtilizationYansong Sun, Jialuo He, Dirk Kutscher, Huangxun ChenMobiCom 2025 · 被引用 1 次
- UDC: Unified DNAS for Compressible TinyML Models for Neural Processing UnitsIgor Fedorov, Ramon Matas Navarro, Hokchhay Tann, Chuteng Zhou 等NeurIPS 2022 · 被引用 19 次
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 被引用 150 次
- Neuro-C: Neural Inference Shaped by Hardware LimitsDiletta Romano, Luca Mottola, Thiemo VoigtEuroSys 2026 · 被引用 2 次
- QUTE: Quantifying Uncertainty in TinyML models with Early-exit-assisted ensembles for model-monitoringNikhil Pratap Ghanathe, Steven J. E. WiltonICML 2025
