Shfl-BW: accelerating deep neural network inference with tensor-core aware weight pruning
Guyue Huang, Haoran Li, Minghai Qin, Fei Sun, Yufei Ding, Yuan Xie
Abstract
Weight pruning in deep neural networks (DNNs) can reduce storage and computation cost, but struggles to bring practical speedup to the model inference time. Tensor-cores can significantly boost the throughput of GPUs on dense computation, but exploiting tensor-cores for sparse DNNs is very challenging. Compared to existing CUDA-cores, tensor-cores require higher data reuse and matrix-shaped instruction granularity, both difficult to yield from sparse DNN kernels. Existing pruning approaches fail to balance the demands of accuracy and efficiency: random sparsity preserves the model quality well but prohibits tensor-core acceleration, while highly-structured block-wise sparsity can exploit tensor-cores but suffers from severe accuracy loss.
In this work, we propose a novel sparse pattern, Shuffled Blockwise sparsity (Shfl-BW ), designed to efficiently utilize tensor-cores while minimizing the constraints on the weight structure. Our insight is that row-and column-wise permutation provides abundant flexibility for the weight structure, while introduces negligible overheads using our GPU kernel designs. We optimize the GPU kernels for Shfl-BW in linear and convolution layers. Evaluations show that our techniques can achieve the state-of-the-art speed-accuracy trade-offs on GPUs. For example, with small accuracy loss, we can accelerate the computation-intensive layers of Transformer [1] by 1.81, 4.18 and 1.90× on NVIDIA V100, T4 and A100 GPUs respectively at 75% sparsity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8407b66a-eec7-42a6-916a-ad938cc9d133Cited by top-tier papers7
- Flash-LLM: Enabling Low-Cost and Highly-Efficient Large Generative Model Inference With Unstructured SparsityHaojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang et al.VLDB 2024 · 29 citations
- Exposing and Exploiting Fine-Grained Block Structures for Fast and Accurate Sparse TrainingPeng Jiang, Lihan Hu, Shihui SongNeurIPS 2022 · 21 citations
- Register Tiling for Unstructured Sparsity in Neural Network InferenceLucas Wilkinson, Kazem Cheshmi, Maryam Mehri DehnaviPLDI 2023 · 17 citations
- TB-STC: Transposable Block-wise N: M Structured Sparse Tensor CoreJun Liu, Shulin Zeng, Junbo Zhao, Li Ding et al.HPCA 2025 · 9 citations
- Harnessing Manycore Processors with Distributed Memory for Accelerated Training of Sparse and Recurrent ModelsJan Finkbeiner, Thomas Gmeinder, Mark Pupilli, Alexander Titterton et al.AAAI 2024 · 7 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- Accelerating sparse DNN models without hardware-support via tile-wise sparsityCong Guo, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu et al.SC 2020 · 65 citations
- Efficient tensor core-based GPU kernels for structured sparsity under reduced precisionZhaodong Chen, Zheng Qu, Liu Liu, Yufei Ding et al.SC 2021 · 54 citations
- Effective Model Sparsification by Scheduled Grow-and-Prune MethodsXiaolong Ma, Minghai Qin, Fei Sun, Zejiang Hou et al.ICLR 2022 · 45 citations
Related papers
- Taming Unstructured Sparsity on GPUs via Latency-Aware OptimizationMaohua Zhu, Yuan XieDAC 2020 · 3 citations
- SlideSparse: Fast and Flexible (2N-2):2N Structured SparsityYingbo HAO, Hanyong Shao, Ting Song, Yan Xia et al.ICML 2026
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo et al.ISCA 2021 · 109 citations
- Eureka: Efficient Tensor Cores for One-sided Unstructured Sparsity in DNN InferenceAshish Gondimalla, Mithuna Thottethodi, T. N. VijaykumarMICRO 2023 · 15 citations
- RM-STC: Row-Merge Dataflow Inspired GPU Sparse Tensor Core for Energy-Efficient Sparse AccelerationGuyue Huang, Zhengyang Wang, Po-An Tsai, Chen Zhang et al.MICRO 2023 · 15 citations
