GPUs All Grown-Up: Fully Device-Driven SpMV Using GPU Work Graphs
Fabian Wildgrube, Pete Ehrett, Paul Trojahn, Richard Membarth, Bradford M. Beckmann, Dominik Baumeister, Matthäus G. Chajdas
摘要
Sparse matrix-vector multiplication (SpMV) is a key operation across high-performance computing, graph analytics, and many more applications.In these applications, the matrix characteristics, notably non-zero elements per row, can vary widely and impact which algorithm performs best.Thus, Graphics Processing Unit (GPU) SpMV algorithms often rely on costly preprocessing to determine what per-row algorithm to select to achieve high performance.In this work we combine SpMV preprocessing and the subsequent per-row processing on the GPU by leveraging the novel "Work Graphs" GPU programming model-initially designed for graphics applications-for dynamic on-device self-scheduling.Work Graphs allow for fine-grain dataflow execution of individual workgroups using emerging hardware and firmware support.As soon as preprocessing has generated sufficient work, workgroups of individual processing kernels are self-scheduled and executed, interleaved with those of other kernels.This improves cache locality and eliminates host interaction altogether.Across a suite of 59 sparse matrices, the best of various novel Work Graphs SpMV implementations outperforms state-of-the-art rocSPARSE "LRB" for a single SpMV by up to 7.19× (mean: 3.35×, SD: 1.89).Furthermore, it achieves much more stable performance across various sparsity patterns than the rocSPARSE CSR-General algorithm, and even beats the advanced rocSPARSE CSR-Adaptive algorithm for up to 92 consecutive SpMV calculations.In addition, compared to rocSPARSE LRB, it reduces code complexity by 75%.Its memory footprint for supporting data structures is a fixed ∼25 MiB independent of matrix size, compared to rocSPARSE LRB's data structures that scale with matrix
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Efficient Algorithm Design of Optimizing SpMV on GPUGenshen Chu, Yuanjie He, Lingyu Dong, Zhezhao Ding 等HPDC 2023 · 被引用 17 次
- spECK: accelerating GPU sparse matrix-matrix multiplication through lightweight analysisMathias Parger, Martin Winter, Daniel Mlakar, Markus SteinbergerPPoPP 2020 · 被引用 48 次
- DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector MultiplicationYuechen Lu, Weifeng LiuSC 2023 · 被引用 37 次
- AlphaSparse: Generating High Performance SpMV Codes Directly from Sparse MatricesZhen Du, Jiajia Li, Yinshan Wang, Xueqi Li 等SC 2022 · 被引用 43 次
- Accelerating SpMV for Scale-Free Graphs with Optimized BinsYuAng Chen, Jeffrey Xu YuICDE 2024 · 被引用 3 次
