Heuristic adaptability to input dynamics for SpMM on CPUs
Guohao Dai, Guyue Huang, Shang Yang, Zhongming Yu, Hengrui Zhang, Yufei Ding, Yuan Xie, Huazhong Yang, Yu Wang
Abstract
Sparse Matrix-Matrix Multiplication (SpMM) has served as fundamental components in various domains. Many previous studies exploit GPUs for SpMM acceleration because GPUs provide high bandwidth and parallelism. We point out that a static design does not always improve the performance of SpMM on different input data (e.g., >85% performance loss with a single algorithm). In this paper, we consider the challenge of input dynamics from a novel autotuning perspective, while following issues remain to be solved: (1) Orthogonal design principles considering sparsity. Orthogonal design principles for such a sparse problem should be extracted to form different algorithms, and further used for performance tuning. (2) Nontrivial implementations in the algorithm space. Combining orthogonal design principles to create new algorithms needs to tackle with new challenges like thread race handling. (3) Heuristic adaptability to input dynamics. The heuristic adaptability is required to dynamically optimize code for input dynamics.
To tackle these challenges, we first propose a novel three-loop model to extract orthogonal design principles for SpMM on GPUs. The model not only covers previous SpMM designs, but also comes up with new designs absent from previous studies. We propose techniques like conditional reduction to implement algorithms missing in previous studies. We further propose DA-SpMM, a Data-Aware heuristic GPU kernel for SpMM. DA-SpMM adaptively optimizes code considering input dynamics. Extensive experimental results show that, DA-SpMM achieves 1.26×∼1.37× speedup compared with the best NVIDIA cuSPARSE algorithm on average, and brings up to 5.59× end-to-end speedup to applications like Graph Neural Networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fa4daa1-9281-46de-8a52-4cfc8d2f2746Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- GE-SpMM: general-purpose sparse matrix-matrix multiplication on GPUs for graph neural networksGuyue Huang, Guohao Dai, Yu Wang, Huazhong YangSC 2020 · 130 citations
- Structured Pruning of Large Language ModelsZiheng Wang, Jeremy Wohlwend, Tao LeiEMNLP 2020 · 88 citations
- FAFNIR: Accelerating Sparse Gathering by Using Efficient Near-Memory Intelligent ReductionBahar Asgari, Ramyad Hadidi, Jiashen Cao, Da Eun Shim et al.HPCA 2021 · 87 citations
Related papers
- HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU CoresZhonggen Li, Xiangyu Ke, Yifan Zhu, Yunjun Gao et al.ICDE 2025 · 5 citations
- Rethinking Tiling and Dataflow for SpMM Acceleration: A Graph Transformation FrameworkAmir Ghazizadeh Ahsaei, Lingxiang Yin, Shilin Tian, Fangzhou Ye et al.MICRO 2025 · 4 citations
- DySpMM: From Fix to Dynamic for Sparse Matrix-Matrix Multiplication AcceleratorsHongyi Wang, Kai Zhong, Haoyu Zhang, Shulin Zeng et al.DAC 2024 · 3 citations
- DTC-SpMM: Bridging the Gap in Accelerating General Sparse Matrix Multiplication with Tensor CoresRuibo Fan, Wei Wang, Xiaowen ChuASPLOS 2024 · 46 citations
- ASM-SpMM: Unleashing the Potential of Arm SME for Sparse Matrix Multiplication AccelerationJiazhi Jiang, Xijia Yao, Jiayu Chen, Jinhui Wei et al.PPoPP 2026
