SC2024Top-tier venue
High Performance Unstructured SpMM Computation Using Tensor Cores
Patrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta, Flavio Vella, Torsten Hoefler
Abstract
High-performance sparse matrix-matrix (SpMM) multiplication is paramount for science and industry, as the ever-increasing sizes of data prohibit using dense data structures. Yet, existing hardware, such as Tensor Cores (TC), is ill-suited for SpMM, as it imposes strict constraints on data structures that cannot be met by unstructured sparsity found in many applications. To address this, we introduce (S)parse (Ma)trix Matrix (T)ensor Core-accelerated (SMaT): a novel SpMM library that utilizes TCs for unstructured sparse matrices. Our block-sparse library leverages the low-level CUDA MMA (matrix-matrix-accumulate) API, maximizing the performance offered by modern GPUs. Algorithmic optimizations such as sparse matrix permutation, further improve performance by minimizing the number of non-zero blocks. The evaluation on NVIDIA A100 shows that SMaT outperforms SotA libraries (DASP, cuSPARSE, and Magicube) by up to 125x (on average 2.6x). SMaT can be used to accelerate many workloads in scientific computing, large model training, inference, and others.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8fdcb7f4-7b74-43af-bcc6-56f46a3e87a2Cited by top-tier papers3
- Characterizing Matrix Multiplication Units across General Parallel Patterns in Scientific ComputingYuechen Lu, Hongwei Zeng, Marc Casas, Weifeng LiuPPoPP 2026 · 1 citation
- Unified Static-Dynamic Pruning for Efficient LLM InferenceJinhyeok Kim, Yejoon Lee, Jaeyoung DoVLDB 2026
- ThunderGNN: Unlocking Tensor Cores for Graph Neural NetworksYuAng Chen, Siyi Teng, Wenqi Weng, Jeffrey Xu YuVLDB 2026
Builds on7
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- Efficient tensor core-based GPU kernels for structured sparsity under reduced precisionZhaodong Chen, Zheng Qu, Liu Liu, Yufei Ding et al.SC 2021 · 54 citations
- AlphaSparse: Generating High Performance SpMV Codes Directly from Sparse MatricesZhen Du, Jiajia Li, Yinshan Wang, Xueqi Li et al.SC 2022 · 43 citations
- DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector MultiplicationYuechen Lu, Weifeng LiuSC 2023 · 37 citations
- GraphMineSuite: Enabling High-Performance and Programmable Graph Mining Algorithms with Set AlgebraMaciej Besta, Zur Vonarburg-Shmaria, Yannick Schaffner, Leonardo Schwarz et al.VLDB 2021 · 28 citations
Related papers
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
- Bridging the Gap between Unstructured SpMM and Structured Sparse Tensor CoresYukang Dong, Ziyuan Shen, Wenbin Jiang, Zhenghang Liu et al.SC 2025 · 4 citations
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
- DTC-SpMM: Bridging the Gap in Accelerating General Sparse Matrix Multiplication with Tensor CoresRuibo Fan, Wei Wang, Xiaowen ChuASPLOS 2024 · 46 citations
- Efficient Quantized Sparse Matrix Operations on Tensor CoresShigang Li, Kazuki Osawa, Torsten HoeflerSC 2022 · 27 citations
