PCNN: Pattern-based Fine-Grained Regular Pruning Towards Optimizing CNN Accelerators
Zhanhong Tan, Jiebo Song, Xiaolong Ma, Sia Huat Tan, Hongyang Chen, Yuanqing Miao, Yifu Wu, Shaokai Ye, Yanzhi Wang, Dehui Li, Kaisheng Ma
Abstract
Weight pruning is a powerful technique to realize model compression. We propose PCNN, a fine-grained regular 1D pruning method. A novel index format called Sparsity Pattern Mask (SPM) is presented to encode the sparsity in PCNN. Leveraging SPM with limited pruning patterns and non-zero sequences with equal length, PCNN can be efficiently employed in hardware. Evaluated on VGG-16 and ResNet-18, our PCNN achieves the compression rate up to 8.4× with only 0.2% accuracy loss. We also implement a pattern-aware architecture in 55nm process, achieving up to 9.0× speedup and 28.39 TOPS/W efficiency with only 3.1% on-chip memory overhead of indices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aea2c59e-63a1-4a7c-b207-638d1a776a04Cited by top-tier papers5
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu et al.ICLR 2021 · 301 citations
- Spada: Accelerating Sparse Matrix Multiplication with Adaptive DataflowZhiyao Li, Jiaxiang Li, Taijie Chen, Dimin Niu et al.ASPLOS 2023 · 59 citations
- Adyna: Accelerating Dynamic Neural Networks with Adaptive SchedulingZhiyao Li, Bohan Yang, Jiaxiang Li, Taijie Chen et al.HPCA 2025 · 2 citations
- SAS: Structured Activation SparsificationYusuke Sekikawa, Shingo YashimaICLR 2024 · 1 citation
- Light-weight Calibrator: A Separable Component for Unsupervised Domain AdaptationShaokai Ye, Kailu Wu, Mu Zhou, Yunfei Yang et al.CVPR 2020
Related papers
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin et al.AAAI 2020 · 201 citations
- PENNI: Pruned Kernel Sharing for Efficient CNN InferenceShiyu Li, Edward Hanson, Hai Li, Yiran ChenICML 2020 · 23 citations
- Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based PermutationXizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying TsuiDAC 2020 · 30 citations
- PIM-Prune: Fine-Grain DCNN Pruning for Crossbar-Based Process-In-Memory ArchitectureChaoqun Chu, Yanzhi Wang, Yilong Zhao, Xiaolong Ma et al.DAC 2020 · 64 citations
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang et al.ASPLOS 2020 · 214 citations
