Accelerating Sparse Convolution with Column Vector-Wise Sparsity
Yijun Tan, Kai Han, Kang Zhao, Xianzhi Yu, Zidong Du, Yunji Chen, Yunhe Wang, Jun Yao
摘要
Weight sparsity is a promising approach to reducing the model size and computation cost of convolutional neural networks (CNNs). Nevertheless, non-zero weights often distribute randomly in sparse CNN models, introducing enormous difficulty in obtaining actual speedup on common hardware (e.g., GPU) over their dense counterparts. Existing acceleration solutions either require hardware modifications for irregular memory access support or rely on a partially structured sparsity pattern. Neither of these methods is capable of achieving fruitful speedup on convolution layers. In this work, we propose an algorithm-software co-designed sparse convolution based on a novel out-vector-wise (OVW) sparse pattern. Building on the insight that vertical vector integrity can preserve continuous memory access in IM2COL, the OVW pattern treats a V × 1 vector as unit. To reduce the error caused by sparsity, we propose an equivalent transformation process, i.e., clustering-based channel permutation, to gather similar rows together. Experimental evaluations demonstrate that our method achieves a 1 . 7 × and 3 . 2 × speedup over the SOTA solution and the dense convolution of ResNet50 on NVIDIA V100 at 75% sparsity, respectively, with only negligible accuracy loss. Moreover, compared to the SOTA solution that achieves speedups only on data with 60% sparsity or more, our method begins to obtain speedups on data with only 10% sparsity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HighLight: Efficient and Flexible DNN Acceleration with Hierarchical Structured SparsityYannan Nellie Wu, Po-An Tsai, Saurav Muralidharan, Angshuman Parashar 等MICRO 2023 · 被引用 29 次
- FlexHiNM-GP: Flexible Hierarchical Pruning via Region Allocation and Channel PermutationXiaodie Yi, Hayun Lee, Dongkun ShinICLR 2026
它引用的顶会 Paper5
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu 等ICLR 2021 · 被引用 301 次
- CHIP: CHannel Independence-based Pruning for Compact Neural NetworksYang Sui, Miao Yin, Yi Xie, Huy Phan 等NeurIPS 2021 · 被引用 198 次
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 被引用 170 次
- Channel Permutations for N: M SparsityJeff Pool, Chong YuNeurIPS 2021 · 被引用 75 次
- Accelerating sparse DNN models without hardware-support via tile-wise sparsityCong Guo, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu 等SC 2020 · 被引用 65 次
相关 Paper
- SUBP: Soft Uniform Block Pruning for 1×N Sparse CNNs Multithreading AccelerationJingyang Xiang, Siqi Li, Jun Chen, Guang Dai 等NeurIPS 2023 · 被引用 2 次
- Sparse Weight Activation TrainingMd Aamir Raihan, Tor M. AamodtNeurIPS 2020 · 被引用 83 次
- Dynamic Sparsity Is Channel-Level Sparsity LearnerLu Yin, Gen Li, Meng Fang, Li Shen 等NeurIPS 2023 · 被引用 29 次
- Efficient tensor core-based GPU kernels for structured sparsity under reduced precisionZhaodong Chen, Zheng Qu, Liu Liu, Yufei Ding 等SC 2021 · 被引用 54 次
- ESCALATE: Boosting the Efficiency of Sparse CNN Accelerator with Kernel DecompositionShiyu Li, Edward Hanson, Xuehai Qian, Hai (Helen) Li 等MICRO 2021 · 被引用 29 次
