Accelerating Sparse Convolution with Column Vector-Wise Sparsity
Yijun Tan, Kai Han, Kang Zhao, Xianzhi Yu, Zidong Du, Yunji Chen, Yunhe Wang, Jun Yao
Abstract
Weight sparsity is a promising approach to reducing the model size and computation cost of convolutional neural networks (CNNs). Nevertheless, non-zero weights often distribute randomly in sparse CNN models, introducing enormous difficulty in obtaining actual speedup on common hardware (e.g., GPU) over their dense counterparts. Existing acceleration solutions either require hardware modifications for irregular memory access support or rely on a partially structured sparsity pattern. Neither of these methods is capable of achieving fruitful speedup on convolution layers. In this work, we propose an algorithm-software co-designed sparse convolution based on a novel out-vector-wise (OVW) sparse pattern. Building on the insight that vertical vector integrity can preserve continuous memory access in IM2COL, the OVW pattern treats a V × 1 vector as unit. To reduce the error caused by sparsity, we propose an equivalent transformation process, i.e., clustering-based channel permutation, to gather similar rows together. Experimental evaluations demonstrate that our method achieves a 1 . 7 × and 3 . 2 × speedup over the SOTA solution and the dense convolution of ResNet50 on NVIDIA V100 at 75% sparsity, respectively, with only negligible accuracy loss. Moreover, compared to the SOTA solution that achieves speedups only on data with 60% sparsity or more, our method begins to obtain speedups on data with only 10% sparsity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- HighLight: Efficient and Flexible DNN Acceleration with Hierarchical Structured SparsityYannan Nellie Wu, Po-An Tsai, Saurav Muralidharan, Angshuman Parashar et al.MICRO 2023 · 29 citations
- FlexHiNM-GP: Flexible Hierarchical Pruning via Region Allocation and Channel PermutationXiaodie Yi, Hayun Lee, Dongkun ShinICLR 2026
Builds on5
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu et al.ICLR 2021 · 301 citations
- CHIP: CHannel Independence-based Pruning for Compact Neural NetworksYang Sui, Miao Yin, Yi Xie, Huy Phan et al.NeurIPS 2021 · 198 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- Channel Permutations for N: M SparsityJeff Pool, Chong YuNeurIPS 2021 · 75 citations
- Accelerating sparse DNN models without hardware-support via tile-wise sparsityCong Guo, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu et al.SC 2020 · 65 citations
Related papers
- SUBP: Soft Uniform Block Pruning for 1×N Sparse CNNs Multithreading AccelerationJingyang Xiang, Siqi Li, Jun Chen, Guang Dai et al.NeurIPS 2023 · 2 citations
- Sparse Weight Activation TrainingMd Aamir Raihan, Tor M. AamodtNeurIPS 2020 · 83 citations
- Dynamic Sparsity Is Channel-Level Sparsity LearnerLu Yin, Gen Li, Meng Fang, Li Shen et al.NeurIPS 2023 · 29 citations
- Efficient tensor core-based GPU kernels for structured sparsity under reduced precisionZhaodong Chen, Zheng Qu, Liu Liu, Yufei Ding et al.SC 2021 · 54 citations
- ESCALATE: Boosting the Efficiency of Sparse CNN Accelerator with Kernel DecompositionShiyu Li, Edward Hanson, Xuehai Qian, Hai (Helen) Li et al.MICRO 2021 · 29 citations
