ZeroBN: Learning Compact Neural Networks For Latency-Critical Edge Systems
Shuo Huai, Lei Zhang, Di Liu, Weichen Liu, Ravi Subramaniam
摘要
Edge devices have been widely adopted to bring deep learning applications onto low power embedded systems, mitigating the privacy and latency issues of accessing cloud servers. The increasingly computational demand of complex neural network models leads to large latency on edge devices with limited resources. Many application scenarios are real-time and have a strict latency constraint, while conventional neural network compression methods are not latency-oriented. In this work, we propose a novel compact neural networks training method to reduce the model latency on latency-critical edge systems. A latency predictor is also introduced to guide and optimize this procedure. Coupled with the latency predictor, our method can guarantee the latency for a compact model by only one training process. The experiment results show that, compared to state-of-the-art model compression methods, our approach can well-fit the ‘hard’ latency constraint by significantly reducing the latency with a mild accuracy drop. To satisfy a 34ms latency constraint, we compact ResNet-50 with 0.82% of accuracy drop. And for GoogLeNet, we can even increase the accuracy by 0.3%
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim 等ISCA 2024 · 被引用 13 次
- Towards Efficient Convolutional Neural Network for Embedded Hardware via Multi-Dimensional PruningHao Kong, Di Liu, Xiangzhong Luo, Shuo Huai 等DAC 2023 · 被引用 2 次
相关 Paper
- Efficient Edge Inference by Selective QueryAnil Kag, Igor Fedorov, Aditya Gangrade, Paul N. Whatmough 等ICLR 2023
- MST-compression: Compressing and Accelerating Binary Neural Networks with Minimum Spanning TreeQuang Hieu Vo, Linh-Tam Tran, Sung-Ho Bae, Lok-Won Kim 等ICCV 2023 · 被引用 2 次
- Latency-aware Spatial-wise Dynamic NetworksYizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue 等NeurIPS 2022 · 被引用 30 次
- Rethinking Pruning for Accelerating Deep Inference At the EdgeDawei Gao, Xiaoxi He, Zimu Zhou, Yongxin Tong 等KDD 2020 · 被引用 24 次
- AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge EnvironmentsHao Wen, Yuanchun Li, Zunshuai Zhang, Shiqi Jiang 等MobiCom 2023 · 被引用 55 次
