Boosting Deep Neural Network Efficiency with Dual-Module Inference
Liu Liu, Lei Deng, Zhaodong Chen, Yuke Wang, Shuangchen Li, Jingwei Zhang, Yihua Yang, Zhenyu Gu, Yufei Ding, Yuan Xie
摘要
Using deep neural networks (DNNs) in machine learning tasks is promising in delivering highquality results but challenging to meet stringent latency requirements and energy constraints because of the memory-bound and the computebound execution pattern of DNNs. We propose a big-little dual-module inference to dynamically skip unnecessary memory accesses and computations to accelerate DNN inference. Leveraging the noise-resilient feature of nonlinear activation functions, we propose to use a lightweight little module that approximates the original DNN layer, termed as the big module, to compute activations of the insensitive region that are more noise-resilient. Hence, the expensive memory accesses and computations of the big module can be reduced as the results are only calculated in the sensitive region. For memory-bound models such as recurrent neural networks (RNNs), our method can reduce the overall memory accesses by 40% on average and achieve 1.54x to 1.75x speedup on a commodity CPU-based server platform with a negligible impact on model quality. In addition, our method can reduce the operations of the compute-bound models such as convolutional neural networks (CNNs) by 3.02x, with only a 0.5% accuracy drop.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Prior Gradient Mask Guided Pruning-Aware Fine-TuningLinhang Cai, Zhulin An, Chuanguang Yang, Yangchun Yan 等AAAI 2022 · 被引用 44 次
- APNN-TC: accelerating arbitrary precision neural networks on ampere GPU tensor coresBoyuan Feng, Yuke Wang, Tong Geng, Ang Li 等SC 2021 · 被引用 36 次
- DUET: Boosting Deep Neural Network Efficiency on Dual-Module ArchitectureLiu Liu, Zheng Qu, Lei Deng, Fengbin Tu 等MICRO 2020 · 被引用 27 次
- TIPS: Topologically Important Path Sampling for Anytime Neural NetworksGuihong Li, Kartikeya Bhardwaj, Yuedong Yang, Radu MarculescuICML 2023
相关 Paper
- Towards Memory-Efficient Neural Networks via Multi-Level in situ GenerationJiaqi Gu, Hanqing Zhu, Chenghao Feng, Mingjie Liu 等ICCV 2021 · 被引用 4 次
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo 等ISCA 2021 · 被引用 109 次
- DRQ: Dynamic Region-based Quantization for Deep Neural Network AccelerationZhuoran Song, Bangqi Fu, Feiyang Wu, Zhaoming Jiang 等ISCA 2020 · 被引用 92 次
- ALERT: Accurate Learning for Energy and TimelinessChengcheng Wan, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann 等USENIX ATC 2020 · 被引用 15 次
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu 等ASPLOS 2026
