Differentiable Neural Network Pruning to Enable Smart Applications on Microcontrollers
Edgar Liberis, Nicholas D. Lane
摘要
Wearable, embedded, and IoT devices are a centrepiece of many ubiquitous computing applications, such as fitness tracking, health monitoring, home security and voice assistants. By gathering user data through a variety of sensors and leveraging machine learning (ML), applications can adapt their behaviour: in other words, devices become "smart". Such devices are typically powered by microcontroller units (MCUs). As MCUs continue to improve, smart devices become capable of performing a non-trivial amount of sensing and data processing, including machine learning inference, which results in a greater degree of user data privacy and autonomy, compared to offloading the execution of ML models to another device. Advanced predictive capabilities across many tasks make neural networks an attractive ML model for ubiquitous computing applications; however, on-device inference on MCUs remains extremely challenging. Orders of magnitude less storage, memory and computational ability, compared to what is typically required to execute neural networks, impose strict structural constraints on the network architecture and call for specialist model compression methodology. In this work, we present a differentiable structured pruning method for convolutional neural networks, which integrates a model's MCU-specific resource usage and parameter importance feedback to obtain highly compressed yet accurate models. Compared to related network pruning work, compressed models are more accurate due to better use of MCU resource budget, and compared to MCU specialist work, compressed models are produced faster. The user only needs to specify the amount of available computational resources and the pruning algorithm will automatically compress the network during training to satisfy them. We evaluate our methodology using benchmark image and audio classification tasks and find that it (a) improves key resource usage of neural networks up to 80x; (b) has little to no overhead or even improves model training time; (c) produces compressed models with matching or improved resource usage up to 1.4x in less time compared to prior MCU-specific model compression methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the EdgeChunlin Tian, Xinpeng Qin, Kahou Tam, Li Li 等USENIX ATC 2025 · 被引用 41 次
- Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning ExperiencesFred Hohman, Mary Beth Kery, Donghao Ren, Dominik MoritzCHI 2024 · 被引用 27 次
- DEX: Data Channel Extension for Efficient CNN Inference on Tiny AI AcceleratorsTaesik Gong, Fahim Kawsar, Chulhong MinNeurIPS 2024 · 被引用 8 次
- Test-Time Adaptation with Binary FeedbackTaeckyung Lee, Sorn Chottananurak, Junsu Kim, Jinwoo Shin 等ICML 2025
它引用的顶会 Paper7
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn 等NeurIPS 2020 · 被引用 827 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- BRP-NAS: Prediction-based NAS using GCNsLukasz Dudziak, Thomas Chau, Mohamed S. Abdelfattah, Royson Lee 等NeurIPS 2020 · 被引用 233 次
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev 等ICLR 2020 · 被引用 229 次
- AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression RatesNing Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang 等AAAI 2020 · 被引用 204 次
相关 Paper
- DTMM: Deploying TinyML Models on Extremely Weak IoT Devices with PruningLixiang Han, Zhen Xiao, Zhenjiang LiINFOCOM 2024 · 被引用 20 次
- MyML: User-Driven Machine LearningVidushi Goyal, Valeria Bertacco, Reetuparna DasDAC 2021 · 被引用 3 次
- UDC: Unified DNAS for Compressible TinyML Models for Neural Processing UnitsIgor Fedorov, Ramon Matas Navarro, Hokchhay Tann, Chuteng Zhou 等NeurIPS 2022 · 被引用 19 次
- Intermittent-Aware Neural Network PruningChih-Chia Lin, Chia-Yin Liu, Chih-Hsuan Yen, Tei-Wei Kuo 等DAC 2023 · 被引用 11 次
- Harmonious Coexistence of Structured Weight Pruning and Ternarization for Deep Neural NetworksLi Yang, Zhezhi He, Deliang FanAAAI 2020 · 被引用 28 次
