MCUNet: Tiny Deep Learning on IoT Devices
Ji Lin, Wei-Ming Chen, Yujun Lin, John Cohn, Chuang Gan, Song Han
Abstract
Machine learning on tiny IoT devices based on microcontroller units (MCU) is appealing but challenging: the memory of microcontrollers is 2-3 orders of magnitude smaller even than mobile phones. We propose MCUNet, a framework that jointly designs the efficient neural architecture (TinyNAS) and the lightweight inference engine (TinyEngine), enabling ImageNet-scale inference on microcontrollers. TinyNAS adopts a two-stage neural architecture search approach that first optimizes the search space to fit the resource constraints, then specializes the network architecture in the optimized search space. TinyNAS can automatically handle diverse constraints (i.e. device, latency, energy, memory) under low search costs. TinyNAS is co-designed with TinyEngine, a memory-efficient inference engine to expand the search space and fit a larger model. TinyEngine adapts the memory scheduling according to the overall network topology rather than layer-wise optimization, reducing the memory usage by 3.4×, and accelerating the inference by 1.7-3.3× compared to TF-Lite Micro [3] and CMSIS-NN [28] . MCUNet is the first to achieves >70% ImageNet top1 accuracy on an off-the-shelf commercial microcontroller, using 3.5× less SRAM and 5.7× less Flash compared to quantized MobileNetV2 and ResNet-18. On visual&audio wake words tasks, MCUNet achieves state-of-the-art accuracy and runs 2.4-3.4× faster than Mo-bileNetV2 and ProxylessNAS-based solutions with 3.7-4.1× smaller peak SRAM. Our study suggests that the era of always-on tiny machine learning on IoT devices has arrived. * Not including the runtime buffer overhead (e.g., Im2Col buffer); the actual memory consumption is larger.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e3b758e-3964-42a8-9c91-78b43346e2c2Cited by top-tier papers53
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang et al.NeurIPS 2022 · 345 citations
- Memory-efficient Patch-based Inference for Tiny Deep LearningJi Lin, Wei-Ming Chen, Han Cai, Chuang Gan et al.NeurIPS 2021 · 190 citations
- Efficient Spatially Sparse Inference for Conditional GANs and Diffusion ModelsMuyang Li, Ji Lin, Chenlin Meng, Stefano Ermon et al.NeurIPS 2022 · 66 citations
- Real-time neural network inference on extremely weak devices: agile offloading with explainable AIKai Huang, Wei GaoMobiCom 2022 · 57 citations
Builds on5
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- On Network Design Spaces for Visual RecognitionIlija Radosavovic, Justin Johnson, Saining Xie, Wan-Yen Lo et al.ICCV 2019 · 148 citations
- Designing Network Design SpacesIlija Radosavovic, Raj Prateek Kosaraju, Ross B. Girshick, Kaiming He et al.CVPR 2020
Related papers
- AtomNet: Designing Tiny Models from Operators Under Extreme MCU ConstraintsZhiwei Dong, Mingzhu Shen, Shihao Bai, Xiuying Wei et al.AAAI 2025
- Entropy-Driven Mixed-Precision Quantization for Deep Network DesignZhenhong Sun, Ce Ge, Junyan Wang, Ming Lin et al.NeurIPS 2022 · 41 citations
- AdaptQNet: Optimizing Quantized DNN on Microcontrollers via Adaptive Heterogeneous Processing Unit UtilizationYansong Sun, Jialuo He, Dirk Kutscher, Huangxun ChenMobiCom 2025 · 1 citation
- MCUFormer: Deploying Vision Tranformers on Microcontrollers with Limited MemoryYinan Liang, Ziwei Wang, Xiuwei Xu, Yansong Tang et al.NeurIPS 2023 · 26 citations
- StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the MicrocontrollerHong-Sheng Zheng, Yu-Yuan Liu, Chen-Fong Hsu, Tsung Tai YehNeurIPS 2023 · 17 citations
