Rethinking Pruning for Accelerating Deep Inference At the Edge
Dawei Gao, Xiaoxi He, Zimu Zhou, Yongxin Tong, Ke Xu, Lothar Thiele
Abstract
There is a growing trend to deploy deep neural networks at the edge for high-accuracy, real-time data mining and user interaction. Applications such as speech recognition and language understanding often apply a deep neural network to encode an input sequence and then use a decoder to generate the output sequence. A promising technique to accelerate these applications on resource-constrained devices is network pruning, which compresses the size of the deep neural network without severe drop in inference accuracy. However, we observe that although existing network pruning algorithms prove effective to speed up the prior deep neural network, they lead to dramatic slowdown of the subsequent decoding and may not always reduce the overall latency of the entire application. To rectify such drawbacks, we propose entropy-based pruning, a new regularizer that can be seamlessly integrated into existing network pruning algorithms. Our key theoretical insight is that reducing the information entropy of the deep neural network outputs decreases the upper bound of the subsequent decoding search space. We validate our solution with two state-of-the-art network pruning algorithms on two model architectures. Experimental results show that compared with existing network pruning algorithms, our entropy-based pruning method notably suppresses and even eliminates the increase of decoding time, and achieves shorter overall latency with only negligible extra accuracy loss in the applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech RecognitionCheng-I Jeff Lai, Yang Zhang, Alexander H. Liu, Shiyu Chang et al.NeurIPS 2021 · 91 citations
- Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech ProcessingYonggan Fu, Yang Zhang, Kaizhi Qian, Zhifan Ye et al.NeurIPS 2022 · 10 citations
- RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language ModelsYufei Li, Zexin Li, Wei Yang, Cong LiuRTSS 2023 · 10 citations
- Pruning-Aware Merging for Efficient Multitask InferenceXiaoxi He, Dawei Gao, Zimu Zhou, Yongxin Tong et al.KDD 2021 · 8 citations
- Context-aware Dynamic Pruning for Speech Foundation ModelsMasao Someki, Yifan Peng, Siddhant Arora, Markus Müller et al.ICLR 2025
Related papers
- Intermittent-Aware Neural Network PruningChih-Chia Lin, Chia-Yin Liu, Chih-Hsuan Yen, Tei-Wei Kuo et al.DAC 2023 · 11 citations
- ZeroBN: Learning Compact Neural Networks For Latency-Critical Edge SystemsShuo Huai, Lei Zhang, Di Liu, Weichen Liu et al.DAC 2021 · 16 citations
- DPACS: Hardware Accelerated Dynamic Neural Network Pruning through Algorithm-Architecture Co-designYizhao Gao, Baoheng Zhang, Xiaojuan Qi, Hayden Kwok-Hay SoASPLOS 2023 · 13 citations
- Differentiable Transportation PruningYunqiang Li, Jan C. van Gemert, Torsten Hoefler, Bert Moons et al.ICCV 2023 · 17 citations
- ALF: Autoencoder-based Low-rank Filter-sharing for Efficient Convolutional Neural NetworksAlexander Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Laura Hauenschild et al.DAC 2020 · 5 citations
