Memory-efficient Patch-based Inference for Tiny Deep Learning
Ji Lin, Wei-Ming Chen, Han Cai, Chuang Gan, Song Han
Abstract
Tiny deep learning on microcontroller units (MCUs) is challenging due to the limited memory size. We find that the memory bottleneck is due to the imbalanced memory distribution in convolutional neural network (CNN) designs: the first several blocks have an order of magnitude larger memory usage than the rest of the network. To alleviate this issue, we propose a generic patch-by-patch inference scheduling, which operates only on a small spatial region of the feature map and significantly cuts down the peak memory. However, naive implementation brings overlapping patches and computation overhead. We further propose receptive field redistribution to shift the receptive field and FLOPs to the later stage and reduce the computation overhead. Manually redistributing the receptive field is difficult. We automate the process with neural architecture search to jointly optimize the neural architecture and inference scheduling, leading to MCUNetV2. Patch-based inference effectively reduces the peak memory usage of existing networks by 4-8×. Co-designed with neural networks, MCUNetV2 sets a record ImageNet accuracy on MCU (71.8%), and achieves >90% accuracy on the visual wake words dataset under only 32kB SRAM. MCUNetV2 also unblocks object detection on tiny devices, achieving 16.9% higher mAP on Pascal VOC compared to the stateof-the-art result. Our study largely addressed the memory bottleneck in tinyML and paved the way for various vision applications beyond image classification. A video demo can be found here. * some CNN designs have highly complicated branching structure (e.g., NASNet [59]), but they are generally less efficient for inference [37, 47, 6] ; thus not widely used for edge computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- FedImpro: Measuring and Improving Client Update in Federated LearningZhenheng Tang, Yonggang Zhang, Shaohuai Shi, Xinmei Tian et al.ICLR 2024 · 25 citations
- Demand Layering for Real-Time DNN Inference with Minimized Memory UsageMingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn et al.RTSS 2022 · 21 citations
- PockEngine: Sparse and Efficient Fine-tuning in a PocketLigeng Zhu, Lanxiang Hu, Ji Lin, Wei-Ming Chen et al.MICRO 2023 · 14 citations
- PROS: an efficient pattern-driven compressive sensing framework for low-power biopotential-based wearables with on-chip intelligenceNhat Pham, Hong Jia, Minh Tran, Tuan Dinh et al.MobiCom 2022 · 12 citations
- Re-thinking computation offload for efficient inference on IoT devices with duty-cycled radiosJin Huang, Hui Guan, Deepak GanesanMobiCom 2023 · 11 citations
Builds on8
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn et al.NeurIPS 2020 · 827 citations
Related papers
- msf-CNN: Patch-based Multi-Stage Fusion with Convolutional Neural Networks for TinyMLZhaolan Huang, Emmanuel BaccelliNeurIPS 2025 · 4 citations
- StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the MicrocontrollerHong-Sheng Zheng, Yu-Yuan Liu, Chen-Fong Hsu, Tsung Tai YehNeurIPS 2023 · 17 citations
- TinyTS: Memory-Efficient TinyML Model Compiler Framework on MicrocontrollersYu-Yuan Liu, Hong-Sheng Zheng, Yu Fang Hu, Chen-Fong Hsu et al.HPCA 2024 · 10 citations
- AtomNet: Designing Tiny Models from Operators Under Extreme MCU ConstraintsZhiwei Dong, Mingzhu Shen, Shihao Bai, Xiuying Wei et al.AAAI 2025
- RaScaNet: Learning Tiny Models by Raster-Scanning ImagesJaehyoung Yoo, Dongwook Lee, Changyong Son, Sangil Jung et al.CVPR 2021
