GUST: Graph Edge-Coloring Utilization for Accelerating Sparse Matrix Vector Multiplication
Armin Gerami, Bahar Asgari
摘要
Sparse matrix-vector multiplication (SpMV) plays a vital role in various scientific and engineering fields, from scientific computing to machine learning. Traditional general-purpose processors often fall short of their peak performance with sparse data, leading to the development of domain-specific architectures to enhance SpMV. Yet, these specialized approaches, whether tailored explicitly for SpMV or adapted from matrix-matrix multiplication accelerators, still face challenges in fully utilizing hardware resources as a result of sparsity. To tackle this problem, we introduce GUST, a hardware/software co-design, the key insight of which lies in separating multipliers and adders in the hardware, thereby enabling resource sharing across multiple rows and columns, leading to efficient hardware utilization and ameliorating negative performance impacts from sparsity. Resource sharing, however, can lead to collisions, a problem we address through a specially devised edge-coloring scheduling algorithm. Our comparisons with various prior domain specific architectures using real-world datasets shows the effectiveness of GUST, with an average hardware utilization of 33.67%. We further evaluate GUST by comparing SpMV execution time and energy consumption of length-256 and -87 GUST with length-256 1-dimensional systolic array (1D), achieving an average speedup of 411× and 108×, and energy efficiency improvement of 137× and 148×, respectively. To asses the implementation aspect, we compare resource consumption of GUST with 1D as a baseline through FPGA synthesis. Length-256 GUST uses the same number of arithmetic units as length-256 1D, while length-87 GUST uses considerably less. We also compare GUST with Serpens, a state-of-the-art FPGA-based SpMV accelerator, with GUST achieving lower execution time on seven out of nine matrices and lower energy consumption on four.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Misam: Machine Learning Assisted Dataflow Selection in Accelerators for Sparse Matrix MultiplicationSanjali Yadav, Amirmahdi Namjoo, Bahar AsgariMICRO 2025 · 被引用 6 次
- Pipirima: Predicting Patterns in Sparsity to Accelerate Matrix AlgebraUbaid Bakhtiar, Donghyeon Joo, Bahar AsgariDAC 2025 · 被引用 6 次
- Chasoň: Supporting Cross HBM Channel Data Migration to Enable Efficient Sparse Algebraic AccelerationUbaid Bakhtiar, Amirmahdi Namjoo, Bahar AsgariMICRO 2025 · 被引用 2 次
- Bootes: Boosting the Efficiency of Sparse Accelerators Using Spectral ClusteringSanjali Yadav, Bahar AsgariMICRO 2025 · 被引用 2 次
它引用的顶会 Paper7
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 被引用 280 次
- Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor ComputationsNitish Kumar Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong 等HPCA 2020 · 被引用 121 次
- SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorXinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak 等HPCA 2021 · 被引用 111 次
- FAFNIR: Accelerating Sparse Gathering by Using Efficient Near-Memory Intelligent ReductionBahar Asgari, Ramyad Hadidi, Jiashen Cao, Da Eun Shim 等HPCA 2021 · 被引用 87 次
- Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignNishil Talati, Kyle May, Armand Behroozi, Yichen Yang 等HPCA 2021 · 被引用 62 次
相关 Paper
- Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationGuowei Zhang, Nithya Attaluri, Joel S. Emer, Daniel SánchezASPLOS 2021 · 被引用 158 次
- SpaHet: A Software/Hardware Co-design for Accelerating Heterogeneous-Sparsity based Sparse Matrix MultiplicationHaoqin Huang, Pengcheng Yao, Zhaozeng An, Yufei Sun 等DAC 2024 · 被引用 3 次
- A Hardware-Software Design Framework for SpMV Acceleration with Flexible Access Pattern PortfolioZhenyu Wu, Maolin Wang, Hayden Kwok-Hay SoHPCA 2025 · 被引用 1 次
- Serpens: a high bandwidth memory based accelerator for general-purpose sparse matrix-vector multiplicationLinghao Song, Yuze Chi, Licheng Guo, Jason CongDAC 2022 · 被引用 56 次
- SegFold: Accelerating Sparse Gemm with a Fine-Grained Dynamic DataflowXinrui Wu, Hanyu Wang, Jason Cong, Tony NowatzkiISCA 2026
