C3ache: Towards Hierarchical Cache-Centric Computing for Sparse Matrix Multiplication on GPGPUs
Xiaojie Li, Mingyu Wang, Baiqing Zhong, Haiqiu Huang, Guangjie Cao, Zhiyi Yu
摘要
Sparse matrix multiplications (SPMMs) are fundamental kernels in various domains and are highly demanded to be executed on general-purpose graphics processing units (GPGPUs).However, it is a challenge to efficiently execute SPMMs across varying sparsity regimes on GPGPUs due to the significant variations in sparsity patterns across different domains.Most of the state-of-the-art works enhance GPGPUs by integrating dedicated accelerator units in processor, which prioritize computational density over memory subsystem optimizations.As sparsity increases, the overhead from expensive memory accesses and redundant input fetches in sparse matrices constrains the effectiveness of processor-centric optimization schemes.Consequently, these optimizations become suboptimal for workloads dominated by irregular memory access patterns and low arithmetic intensity.To address these problems, we proposed a hierarchical cachecentric computing architecture with hybrid dataflow, 𝐶 3 𝑎𝑐ℎ𝑒, to achieve near-optimal memory efficiency and data reuse.First, the hybrid dataflow based on the outer-product dataflow and Gustavson's dataflow is proposed to decouple the SPMM computation into two distinct phases (multiplication and merging) with different behavioral characteristics and align with the memory access pattern.Furthermore, 𝐶 3 𝑎𝑐ℎ𝑒 restructures the cache hierarchy through in-cache computing, transforming the two levels of cache into largescale data parallel processing in memory (PIM) units and in-situ merging PIM units respectively to map the two computation phases of SPMMs.To synchronize the granularity of data fetching with the memory access pattern, a novel memory-aware compressed format is proposed for sparse encoding to further reduce the memory transaction and the decoding overhead of 𝐶 3 𝑎𝑐ℎ𝑒.To realize 𝐶 3 𝑎𝑐ℎ𝑒, an
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 被引用 280 次
- Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationGuowei Zhang, Nithya Attaluri, Joel S. Emer, Daniel SánchezASPLOS 2021 · 被引用 158 次
- ACES: Accelerating Sparse Matrix Multiplication with Adaptive Execution Flow and Concurrency-Aware Cache OptimizationsXiaoyang Lu, Boyu Long, Xiaoming Chen, Yinhe Han 等ASPLOS 2024 · 被引用 13 次
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou 等PPoPP 2025 · 被引用 18 次
- SPAGHETTI: Streaming Accelerators for Highly Sparse GEMM on FPGAsReza Hojabr, Ali Sedaghati, Amirali Sharifian, Ahmad Khonsari 等HPCA 2021 · 被引用 66 次
