MoteNN: Memory Optimization via Fine-grained Scheduling for Deep Neural Networks on Tiny Devices
Renze Chen, Zijian Ding, Size Zheng, Meng Li, Yun Liang
Abstract
There has been a growing trend in deploying deep neural networks (DNNs) on tiny devices. However, deploying DNNs on such devices poses significant challenges due to the contradiction between DNNs' substantial memory requirements and the stringent memory constraints of tiny devices. Some prior works incur large latency overhead to save memory and target only simple CNNs, while others employ coarse-grained scheduling for complicated networks, leading to limited memory footprint reduction. This paper proposes MoteNN that performs fine-grained scheduling via operator partitioning on arbitrary DNNs to dramatically reduce peak memory usage with little latency overhead. MoteNN presents a graph representation named Axis Connecting Graph (ACG) to perform operator partition at graph-level efficiently. MoteNN further proposes an algorithm that finds the partition and schedule guided by memory bottlenecks. We evaluate MoteNN using various popular networks and show that MoteNN achieves up to 80% of peak memory usage reduction compared to the state-of-art works with nearly no latency overhead on tiny devices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69a60402-bf2b-4fb8-b332-016be8685c40Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Exploring Randomly Wired Neural Networks for Image RecognitionSaining Xie, Alexander Kirillov, Ross B. Girshick, Kaiming HeICCV 2019 · 384 citations
- Fast and Practical Neural Architecture SearchJiequan Cui, Pengguang Chen, Ruiyu Li, Shu Liu et al.ICCV 2019 · 69 citations
- Hierarchical memory-constrained operator scheduling of neural architecture search networksZihan Wang, Chengcheng Wan, Yuting Chen, Ziyi Lin et al.DAC 2022 · 16 citations
- You only search once: on lightweight differentiable architecture search for resource-constrained embedded platformsXiangzhong Luo, Di Liu, Hao Kong, Shuo Huai et al.DAC 2022 · 13 citations
- HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML PlatformsJosse Van Delm, Maarten Vandersteegen, Alessio Burrello, Giuseppe Maria Sarda et al.DAC 2023 · 10 citations
Related papers
- FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge DevicesXiangyu Li, Yuanchun Li, Yuanzhe Li, Ting Cao et al.MobiCom 2024 · 40 citations
- Memory-efficient Patch-based Inference for Tiny Deep LearningJi Lin, Wei-Ming Chen, Han Cai, Chuang Gan et al.NeurIPS 2021 · 190 citations
- MAGIS: Memory Optimization via Coordinated Graph Transformation and Scheduling for DNNRenze Chen, Zijian Ding, Size Zheng, Chengrui Zhang et al.ASPLOS 2024 · 14 citations
- TinyTS: Memory-Efficient TinyML Model Compiler Framework on MicrocontrollersYu-Yuan Liu, Hong-Sheng Zheng, Yu Fang Hu, Chen-Fong Hsu et al.HPCA 2024 · 10 citations
- Context-aware Adaptive Surgery: A Fast and Effective Framework for Adaptative Model PartitionHongli Wang, Bin Guo, Jiaqi Liu, Sicong Liu et al.UbiComp 2021 · 25 citations
