DPACS: Hardware Accelerated Dynamic Neural Network Pruning through Algorithm-Architecture Co-design
Yizhao Gao, Baoheng Zhang, Xiaojuan Qi, Hayden Kwok-Hay So
Abstract
By eliminating compute operations intelligently based on the run time input, dynamic pruning (DP) promises to improve deep neural network inference speed substantially without incurring a major impact on their accuracy. Although many DP algorithms with good pruning performance have been proposed, it remains a challenge to translate these theoretical reductions in compute operations into satisfactory end-to-end speedups in practical real-world implementations. The overhead of identifying operations to be pruned during run time, the need to efficiently process the resulting dynamic dataflow, and the non-trivial memory I/O bottleneck that emerges as the number of compute operations reduces, have all contributed to the challenge of implementing practical DP systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 57a02275-4976-48da-98ce-c11bcb3420d9Cited by top-tier papers1
Ask how each one uses itRelated papers
- PowerPruning: Selecting Weights and Activations for Power-Efficient Neural Network AccelerationRichard Petri, Grace Li Zhang, Yiran Chen, Ulf Schlichtmann et al.DAC 2023 · 11 citations
- Rethinking Pruning for Accelerating Deep Inference At the EdgeDawei Gao, Xiaoxi He, Zimu Zhou, Yongxin Tong et al.KDD 2020 · 24 citations
- SPDY: Accurate Pruning with Speedup GuaranteesElias Frantar, Dan AlistarhICML 2022 · 45 citations
- Dynamic Structure Pruning for Compressing CNNsJun-Hyung Park, Yeachan Kim, Junho Kim, Joon-Young Choi et al.AAAI 2023 · 24 citations
- Efficient Latency-Aware CNN Depth Compression via Two-Stage Dynamic ProgrammingJinuk Kim, Yeonwoo Jeong, Deokjae Lee, Hyun Oh SongICML 2023 · 1 citation
