Optimization of GPU-based Sparse Matrix Multiplication for Large Sparse Networks
Jeongmyung Lee, Seokwon Kang, Yongseung Yu, Yong-Yeon Jo, Sang-Wook Kim, Yongjun Park
Abstract
Sparse matrix multiplication (spGEMM) is widely used to analyze the sparse network data, and extract important information based on matrix representation. As it contains a high degree of data parallelism, many efficient implementations using data-parallel programming platforms such as CUDA and OpenCL have been introduced on graphic processing units (GPUs). Several well-known spGEMM techniques, such as cuS- PARSE and CUSP, often do not utilize the GPU resources fully, owing to the load imbalance between threads in the expansion process and high memory contention in the merge process. Furthermore, even though several outer-product-based spGEMM techniques are proposed to solve the load balancing problem on expansion, they still do not utilize the GPU resources fully, because severe computation load variations exist among the multiple thread blocks.To solve these challenges, this paper proposes a new optimization pass called Block Reorganizer, which balances the total computations of each computing unit on target GPUs, based on the outer-product-based expansion process, and reduces the memory pressure during the merge process. For expansion, it first identifies the actual computation amount for each block, and then performs two thread block transformation processes based on their characteristics: 1) B-Splitting to transform a heavy-computation blocks into multiple small blocks and 2) B- Gathering to aggregate multiple small-computation blocks to a larger block. While merging, it improves the overall performance by performing B-Limiting to limit the number of blocks on each computing unit. Experimental results show that it improves the total performance of kernel execution by 1.43x, on an average, when compared to the row-product-based spGEMM, for NVIDIA Titan Xp GPUs on real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1cc1e13-b0a4-4d57-b2cc-6b2d4a37b4c5Cited by top-tier papers3
- TileSpGEMM: a tiled algorithm for parallel sparse general matrix-matrix multiplication on GPUsYuyao Niu, Zhengyang Lu, Haonan Ji, Shuhui Song et al.PPoPP 2022 · 66 citations
- On Efficient Large Sparse Matrix Chain MultiplicationChunxu Lin, Wensheng Luo, Yixiang Fang, Chenhao Ma et al.SIGMOD 2024 · 4 citations
- Going Deeper into Locally Differentially Private Graph Neural NetworksLongzhu He, Chaozhuo Li, Peng Tang, Sen SuICML 2025
Related papers
- Orchestrating Large-Scale SpGEMMs using Dynamic Block Distribution and Data Transfer Minimization on Heterogeneous SystemsTaehyeong Park, Seokwon Kang, Myung-Hwan Jang, Sang-Wook Kim et al.ICDE 2023 · 6 citations
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
- DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector MultiplicationYuechen Lu, Weifeng LiuSC 2023 · 37 citations
- HARP: Hardware-Based Pseudo-Tiling for Sparse Matrix Multiplication AcceleratorJinkwon Kim, Myeongjae Jang, Haejin Nam, Soontae KimMICRO 2023 · 12 citations
