SC2023Top-tier venue
DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector Multiplication
Yuechen Lu, Weifeng Liu
Abstract
Sparse matrix-vector multiplication (SpMV) plays a key role in computational science and engineering, graph processing, and machine learning applications. Much work on SpMV was devoted to resolving problems such as random access to the vector 𝑥 and unbalanced load. However, we have experimentally found that the computation of inner products still occupies much overhead in the SpMV operation, which has been largely ignored in existing work.
In this paper, we propose DASP, a new algorithm using specific dense MMA units for accelerating the compute part of general SpMV. We analyze the row-wise distribution of nonzeros and group the rows into three categories containing long, medium, and short rows, respectively. We then organize them into small blocks of proper sizes to meet the requirement of MMA computation. For the three categories, DASP offers different strategies to complete SpMV by efficiently utilizing the MMA units.
The experimental results on two newest NVIDIA GPUs A100 and H800 show that our DASP in FP64 precision outperforms five latest SpMV methods CSR5, TileSpMV, LSRB-CSR, cuSPARSE BSR format and cuSPARSE CSR format by a factor of on average 1.46x, 2.09x, 3.29x, 2.08x and 1.52x (up to 12.64x, 17.48x, 90.59x, 283.92x and 6.94x) on A100, respectively. As for SpMV in FP16 precision, our DASP outperforms cuSPARSE by a factor of on average 1.70x and 1.75x (up to 26.47x and 65.94x) on A100 and H800, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1f9c8d4-4ff3-40c2-bcbb-1a75b805b978Cited by top-tier papers10
- Mastering Sparse CUDA Generation through Pretrained Models and Deep Reinforcement LearningYaoyu Wang, Hankun Dai, Zhidong Yang, Junmin Xiao et al.ICLR 2026 · 476 citations
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
- AmgT: Algebraic Multigrid Solver on Tensor CoresYuechen Lu, Lijie Zeng, Tengcheng Wang, Xu Fu et al.SC 2024 · 17 citations
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta et al.SC 2024 · 15 citations
- Mille-feuille: A Tile-Grained Mixed Precision Single-Kernel Conjugate Gradient Solver on GPUsDechuang Yang, Yuxuan Zhao, Yiduo Niu, Weile Jia et al.SC 2024 · 8 citations
Builds on10
- TileSpGEMM: a tiled algorithm for parallel sparse general matrix-matrix multiplication on GPUsYuyao Niu, Zhengyang Lu, Haonan Ji, Shuhui Song et al.PPoPP 2022 · 66 citations
- Efficient tensor core-based GPU kernels for structured sparsity under reduced precisionZhaodong Chen, Zheng Qu, Liu Liu, Yufei Ding et al.SC 2021 · 54 citations
- Efficiently running SpMV on long vector architecturesConstantino Gómez, Filippo Mantovani, Erich Focht, Marc CasasPPoPP 2021 · 48 citations
- QGTC: accelerating quantized graph neural networks via GPU tensor coreYuke Wang, Boyuan Feng, Yufei DingPPoPP 2022 · 47 citations
- AlphaSparse: Generating High Performance SpMV Codes Directly from Sparse MatricesZhen Du, Jiajia Li, Yinshan Wang, Xueqi Li et al.SC 2022 · 43 citations
Related papers
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
- Efficient Algorithm Design of Optimizing SpMV on GPUGenshen Chu, Yuanjie He, Lingyu Dong, Zhezhao Ding et al.HPDC 2023 · 17 citations
- Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor CoresKaige Zhang, Hailong Yang, Xin You, Tianyu Feng et al.PPoPP 2026
- Optimization of GPU-based Sparse Matrix Multiplication for Large Sparse NetworksJeongmyung Lee, Seokwon Kang, Yongseung Yu, Yong-Yeon Jo et al.ICDE 2020 · 22 citations
- SpV8: Pursuing Optimal Vectorization and Regular Computation Pattern in SpMVChenyang Li, Tian Xia, Wenzhe Zhao, Nanning Zheng et al.DAC 2021 · 17 citations
