DTC-SpMM: Bridging the Gap in Accelerating General Sparse Matrix Multiplication with Tensor Cores
Ruibo Fan, Wei Wang, Xiaowen Chu
摘要
Sparse Matrix-Matrix Multiplication (SpMM) is a building-block operation in scientific computing and machine learning applications. Recent advancements in hardware, notably Tensor Cores (TCs), have created promising opportunities for accelerating SpMM. However, harnessing these hardware accelerators to speed up general SpMM necessitates considerable effort. In this paper, we undertake a comprehensive analysis of the state-of-the-art techniques for accelerating TC-based SpMM and identify crucial performance gaps. Drawing upon these insights, we propose DTC-SpMM, a novel approach with systematic optimizations tailored for accelerating general SpMM on TCs. DTC-SpMM encapsulates diverse aspects, including efficient compression formats, reordering methods, and runtime pipeline optimizations. Our extensive experiments on modern GPUs with a diverse range of benchmark matrices demonstrate remarkable performance improvements in SpMM acceleration by TCs in conjunction with our proposed optimizations. The case study also shows that DTC-SpMM speeds up end-to-end GNN training by up to 1.91× against popular GNN frameworks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Mastering Sparse CUDA Generation through Pretrained Models and Deep Reinforcement LearningYaoyu Wang, Hankun Dai, Zhidong Yang, Junmin Xiao 等ICLR 2026 · 被引用 476 次
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou 等PPoPP 2025 · 被引用 18 次
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu 等PPoPP 2025 · 被引用 18 次
- AmgT: Algebraic Multigrid Solver on Tensor CoresYuechen Lu, Lijie Zeng, Tengcheng Wang, Xu Fu 等SC 2024 · 被引用 17 次
- Voltrix: Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel OptimizationYaqi Xia, Weihu Wang, Donglin Yang, Xiaobo Zhou 等USENIX ATC 2025 · 被引用 6 次
它引用的顶会 Paper15
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 被引用 170 次
- GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUsYuke Wang, Boyuan Feng, Gushu Li, Shuangchen Li 等OSDI 2021 · 被引用 163 次
- GE-SpMM: general-purpose sparse matrix-matrix multiplication on GPUs for graph neural networksGuyue Huang, Guohao Dai, Yu Wang, Huazhong YangSC 2020 · 被引用 130 次
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo 等ISCA 2021 · 被引用 109 次
相关 Paper
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta 等SC 2024 · 被引用 15 次
- HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU CoresZhonggen Li, Xiangyu Ke, Yifan Zhu, Yunjun Gao 等ICDE 2025 · 被引用 5 次
- Groot: Graph-Centric Row Reordering with Tree for Sparse Matrix Multiplications on Tensor CoresYuAng Chen, Jiadong Xie, Siyi Teng, Wenqi Zeng 等EuroSys 2025 · 被引用 4 次
- Bridging the Gap between Unstructured SpMM and Structured Sparse Tensor CoresYukang Dong, Ziyuan Shen, Wenbin Jiang, Zhenghang Liu 等SC 2025 · 被引用 4 次
- TC-GNN: Bridging Sparse GNN Computation and Dense Tensor Cores on GPUsYuke Wang, Boyuan Feng, Zheng Wang, Guyue Huang 等USENIX ATC 2023
