Boosting Deep Neural Network Efficiency with Dual-Module Inference
Liu Liu, Lei Deng, Zhaodong Chen, Yuke Wang, Shuangchen Li, Jingwei Zhang, Yihua Yang, Zhenyu Gu, Yufei Ding, Yuan Xie
Abstract
Using deep neural networks (DNNs) in machine learning tasks is promising in delivering highquality results but challenging to meet stringent latency requirements and energy constraints because of the memory-bound and the computebound execution pattern of DNNs. We propose a big-little dual-module inference to dynamically skip unnecessary memory accesses and computations to accelerate DNN inference. Leveraging the noise-resilient feature of nonlinear activation functions, we propose to use a lightweight little module that approximates the original DNN layer, termed as the big module, to compute activations of the insensitive region that are more noise-resilient. Hence, the expensive memory accesses and computations of the big module can be reduced as the results are only calculated in the sensitive region. For memory-bound models such as recurrent neural networks (RNNs), our method can reduce the overall memory accesses by 40% on average and achieve 1.54x to 1.75x speedup on a commodity CPU-based server platform with a negligible impact on model quality. In addition, our method can reduce the operations of the compute-bound models such as convolutional neural networks (CNNs) by 3.02x, with only a 0.5% accuracy drop.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bc22d08-c07e-4020-a5c8-d0edb856cae2Cited by top-tier papers4
- Prior Gradient Mask Guided Pruning-Aware Fine-TuningLinhang Cai, Zhulin An, Chuanguang Yang, Yangchun Yan et al.AAAI 2022 · 44 citations
- APNN-TC: accelerating arbitrary precision neural networks on ampere GPU tensor coresBoyuan Feng, Yuke Wang, Tong Geng, Ang Li et al.SC 2021 · 36 citations
- DUET: Boosting Deep Neural Network Efficiency on Dual-Module ArchitectureLiu Liu, Zheng Qu, Lei Deng, Fengbin Tu et al.MICRO 2020 · 27 citations
- TIPS: Topologically Important Path Sampling for Anytime Neural NetworksGuihong Li, Kartikeya Bhardwaj, Yuedong Yang, Radu MarculescuICML 2023
Related papers
- Towards Memory-Efficient Neural Networks via Multi-Level in situ GenerationJiaqi Gu, Hanqing Zhu, Chenghao Feng, Mingjie Liu et al.ICCV 2021 · 4 citations
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo et al.ISCA 2021 · 109 citations
- DRQ: Dynamic Region-based Quantization for Deep Neural Network AccelerationZhuoran Song, Bangqi Fu, Feiyang Wu, Zhaoming Jiang et al.ISCA 2020 · 92 citations
- ALERT: Accurate Learning for Energy and TimelinessChengcheng Wan, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann et al.USENIX ATC 2020 · 15 citations
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu et al.ASPLOS 2026
