SC2023Top-tier venue
VENOM: A Vectorized N: M Format for Unleashing the Power of Sparse Tensor Cores
Roberto L. Castro, Andrei Ivanov, Diego Andrade, Tal Ben-Nun, Basilio B. Fraguela, Torsten Hoefler
Abstract
The increasing success and scaling of Deep Learning models demands higher computational efficiency and power. Sparsification can lead to both smaller models as well as higher compute efficiency, and accelerated hardware is becoming available. However, exploiting it efficiently requires kernel implementations, pruning algorithms, and storage formats, to utilize hardware support of specialized sparse vector units. An example of those are the NVIDIA's Sparse Tensor Cores (SPTCs), which promise a 2× speedup. However, SPTCs only support the 2:4 format, limiting achievable sparsity ratios to 50%. We present the V:N:M format, which enables the execution of arbitrary N:M ratios on SPTCs. To efficiently exploit the resulting format, we propose Spatha, a high-performance sparse-library for DL routines. We show that Spatha achieves up to 37× speedup over cuBLAS. We also demonstrate a second-order pruning technique that enables sparsification to high sparsity ratios with V:N:M and little to no loss in accuracy in modern transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b92de63-c847-42b6-92d1-1074d9ab3ca4Cited by top-tier papers11
- RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust AdaptationMahdi Nikdan, Soroush Tabesh, Elvir Crncevic, Dan AlistarhICML 2024 · 53 citations
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta et al.SC 2024 · 15 citations
- DELTA4: Sparse Matrix-Vector Multiplication for Low SparsityVladimír Macko, Vladimír BožaICML 2026 · 9 citations
- GeneralSparse: Bridging the Gap in SpMM for Pruned Large Language Model Inference on GPUsYaoyu Wang, Xiao Guo, Junmin Xiao, De Chen et al.USENIX ATC 2025 · 5 citations
- Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor CoresChenpeng Wu, Qiqi Gu, Heng Shi, Jianguo Yao et al.EuroSys 2025 · 5 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- PLATON: Pruning Large Transformer Models with Upper Confidence Bound of Weight ImportanceQingru Zhang, Simiao Zuo, Chen Liang, Alexander Bukharin et al.ICML 2022 · 107 citations
Related papers
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
- Efficient Quantized Sparse Matrix Operations on Tensor CoresShigang Li, Kazuki Osawa, Torsten HoeflerSC 2022 · 27 citations
- SPLAT: A Framework for Optimised GPU Code-Generation for SParse reguLar ATtentionAhan Gupta, Yueming Yuan, Devansh Jain, Yuhao Ge et al.OOPSLA 2025 · 3 citations
- Bridging the Gap between Unstructured SpMM and Structured Sparse Tensor CoresYukang Dong, Ziyuan Shen, Wenbin Jiang, Zhenghang Liu et al.SC 2025 · 4 citations
- Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor CoresKaige Zhang, Hailong Yang, Xin You, Tianyu Feng et al.PPoPP 2026
