ZeroBN: Learning Compact Neural Networks For Latency-Critical Edge Systems
Shuo Huai, Lei Zhang, Di Liu, Weichen Liu, Ravi Subramaniam
Abstract
Edge devices have been widely adopted to bring deep learning applications onto low power embedded systems, mitigating the privacy and latency issues of accessing cloud servers. The increasingly computational demand of complex neural network models leads to large latency on edge devices with limited resources. Many application scenarios are real-time and have a strict latency constraint, while conventional neural network compression methods are not latency-oriented. In this work, we propose a novel compact neural networks training method to reduce the model latency on latency-critical edge systems. A latency predictor is also introduced to guide and optimize this procedure. Coupled with the latency predictor, our method can guarantee the latency for a compact model by only one training process. The experiment results show that, compared to state-of-the-art model compression methods, our approach can well-fit the ‘hard’ latency constraint by significantly reducing the latency with a mild accuracy drop. To satisfy a 34ms latency constraint, we compact ResNet-50 with 0.82% of accuracy drop. And for GoogLeNet, we can even increase the accuracy by 0.3%
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 26439cee-633f-4b1d-9aae-4dfdaeec6e66Cited by top-tier papers2
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim et al.ISCA 2024 · 13 citations
- Towards Efficient Convolutional Neural Network for Embedded Hardware via Multi-Dimensional PruningHao Kong, Di Liu, Xiangzhong Luo, Shuo Huai et al.DAC 2023 · 2 citations
Related papers
- Efficient Edge Inference by Selective QueryAnil Kag, Igor Fedorov, Aditya Gangrade, Paul N. Whatmough et al.ICLR 2023
- MST-compression: Compressing and Accelerating Binary Neural Networks with Minimum Spanning TreeQuang Hieu Vo, Linh-Tam Tran, Sung-Ho Bae, Lok-Won Kim et al.ICCV 2023 · 2 citations
- Latency-aware Spatial-wise Dynamic NetworksYizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue et al.NeurIPS 2022 · 30 citations
- Rethinking Pruning for Accelerating Deep Inference At the EdgeDawei Gao, Xiaoxi He, Zimu Zhou, Yongxin Tong et al.KDD 2020 · 24 citations
- AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge EnvironmentsHao Wen, Yuanchun Li, Zunshuai Zhang, Shiqi Jiang et al.MobiCom 2023 · 55 citations
