DTMM: Deploying TinyML Models on Extremely Weak IoT Devices with Pruning
Lixiang Han, Zhen Xiao, Zhenjiang Li
摘要
DTMM is a library designed for efficient deployment and execution of machine learning models on weak IoT devices such as microcontroller units (MCUs). The motivation for designing DTMM comes from the emerging field of tiny machine learning (TinyML), which explores extending the reach of machine learning to many low-end IoT devices to achieve ubiquitous intelligence. Due to the weak capability of embedded devices, it is necessary to compress models by pruning enough weights before deploying. Although pruning has been studied extensively on many computing platforms, two key issues with pruning methods are exacerbated on MCUs: models need to be deeply compressed without significantly compromising accuracy, and they should perform efficiently after pruning. Current solutions only achieve one of these objectives, but not both. In this paper, we find that pruned models have great potential for efficient deployment and execution on MCUs. Therefore, we propose DTMM with pruning unit selection, pre-execution pruning optimizations, runtime acceleration, and post-execution low-cost storage to fill the gap for efficient deployment and execution of pruned models. It can be integrated into commercial ML frameworks for practical deployment, and a prototype system has been developed. Extensive experiments on various models show promising gains compared to state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn 等NeurIPS 2020 · 被引用 827 次
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev 等ICLR 2020 · 被引用 229 次
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang 等ASPLOS 2020 · 被引用 214 次
- CHIP: CHannel Independence-based Pruning for Compact Neural NetworksYang Sui, Miao Yin, Yi Xie, Huy Phan 等NeurIPS 2021 · 被引用 198 次
- Pruning Filter in FilterFanxu Meng, Hao Cheng, Ke Li, Huixiang Luo 等NeurIPS 2020 · 被引用 130 次
相关 Paper
- Differentiable Neural Network Pruning to Enable Smart Applications on MicrocontrollersEdgar Liberis, Nicholas D. LaneUbiComp 2023 · 被引用 27 次
- Intermittent-Aware Neural Network PruningChih-Chia Lin, Chia-Yin Liu, Chih-Hsuan Yen, Tei-Wei Kuo 等DAC 2023 · 被引用 11 次
- TinyTS: Memory-Efficient TinyML Model Compiler Framework on MicrocontrollersYu-Yuan Liu, Hong-Sheng Zheng, Yu Fang Hu, Chen-Fong Hsu 等HPCA 2024 · 被引用 10 次
- UDC: Unified DNAS for Compressible TinyML Models for Neural Processing UnitsIgor Fedorov, Ramon Matas Navarro, Hokchhay Tann, Chuteng Zhou 等NeurIPS 2022 · 被引用 19 次
- Memory-Efficient and Secure DNN Inference on TrustZone-enabled Consumer IoT DevicesXueshuo Xie, Haoxu Wang, Zhaolong Jian, Tao Li 等INFOCOM 2024 · 被引用 11 次
