Taming Unstructured Sparsity on GPUs via Latency-Aware Optimization
Maohua Zhu, Yuan Xie
Abstract
Neural Networks (NNs) exhibit high redundancy in their parameters so that pruning methods can achieve high compression ratio without accuracy loss. However, the very high sparsity produced by unstructured pruning methods is difficult to be efficiently mapped onto Graphics Processing Units (GPUs) because of its decoding overhead and workload imbalance. With the introduction of Tensor Core, the latest GPUs achieve even higher throughput for the dense neural networks. This makes unstructured neural networks fail to outperform their dense counterparts because they are not currently supported by Tensor Core. To tackle this problem, prior work suggests structured pruning to improve the performance of sparse NNs on GPUs. However, such structured pruning methods have to sacrifice a significant part of sparsity to retain the model accuracy, which limits the speedup on the hardware. In this paper, we observe that the Tensor Core is also able to compute unstructured sparse NNs efficiently. To achieve this goal, we first propose ExTensor, a set of sparse Tensor Core instructions with a variable input matrix tile size. The variable tile size allows a matrix multiplication to be implemented by mixing different types of ExTensor instructions. We build a performance model to estimate the latency of an ExTensor instruction given an operand sparse weight matrix. Based on this model, we propose a heuristic algorithm to find the optimal sequence of the instructions for an ExTensor based kernel to achieve the best performance on the GPU. Experimental results demonstrate that our approach achieves 36% better performance than the state-of-the-art sparse Tensor Core design.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 38f7ec7c-2588-4b59-9490-8625e4b5d2cfCited by top-tier papers3
- Flexagon: A Multi-dataflow Sparse-Sparse Matrix Multiplication Accelerator for Efficient DNN ProcessingFrancisco Muñoz-Martínez, Raveesh Garg, Michael Pellauer, José L. Abellán et al.ASPLOS 2023 · 60 citations
- FIARSE: Model-Heterogeneous Federated Learning via Importance-Aware Submodel ExtractionFeijie Wu, Xingchen Wang, Yaqing Wang, Tianci Liu et al.NeurIPS 2024 · 47 citations
- Transformers meet Stochastic Block Models: Attention with Data-Adaptive Sparsity and CostSungjun Cho, Seonwoo Min, Jinwoo Kim, Moontae Lee et al.NeurIPS 2022 · 5 citations
Related papers
- Accelerating sparse DNN models without hardware-support via tile-wise sparsityCong Guo, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu et al.SC 2020 · 65 citations
- Shfl-BW: accelerating deep neural network inference with tensor-core aware weight pruningGuyue Huang, Haoran Li, Minghai Qin, Fei Sun et al.DAC 2022 · 19 citations
- Eureka: Efficient Tensor Cores for One-sided Unstructured Sparsity in DNN InferenceAshish Gondimalla, Mithuna Thottethodi, T. N. VijaykumarMICRO 2023 · 15 citations
- Fractal: Joint Multi-Level Sparse Pattern Tuning of Accuracy and Performance for DNN PruningYue Guan, Changming Yu, Yangjie Zhou, Jingwen Leng et al.ASPLOS 2024 · 6 citations
- Bridging the Gap between Unstructured SpMM and Structured Sparse Tensor CoresYukang Dong, Ziyuan Shen, Wenbin Jiang, Zhenghang Liu et al.SC 2025 · 4 citations
