C3ache: Towards Hierarchical Cache-Centric Computing for Sparse Matrix Multiplication on GPGPUs
Xiaojie Li, Mingyu Wang, Baiqing Zhong, Haiqiu Huang, Guangjie Cao, Zhiyi Yu
Abstract
Sparse matrix multiplications (SPMMs) are fundamental kernels in various domains and are highly demanded to be executed on general-purpose graphics processing units (GPGPUs).However, it is a challenge to efficiently execute SPMMs across varying sparsity regimes on GPGPUs due to the significant variations in sparsity patterns across different domains.Most of the state-of-the-art works enhance GPGPUs by integrating dedicated accelerator units in processor, which prioritize computational density over memory subsystem optimizations.As sparsity increases, the overhead from expensive memory accesses and redundant input fetches in sparse matrices constrains the effectiveness of processor-centric optimization schemes.Consequently, these optimizations become suboptimal for workloads dominated by irregular memory access patterns and low arithmetic intensity.To address these problems, we proposed a hierarchical cachecentric computing architecture with hybrid dataflow, ๐ถ 3 ๐๐โ๐, to achieve near-optimal memory efficiency and data reuse.First, the hybrid dataflow based on the outer-product dataflow and Gustavson's dataflow is proposed to decouple the SPMM computation into two distinct phases (multiplication and merging) with different behavioral characteristics and align with the memory access pattern.Furthermore, ๐ถ 3 ๐๐โ๐ restructures the cache hierarchy through in-cache computing, transforming the two levels of cache into largescale data parallel processing in memory (PIM) units and in-situ merging PIM units respectively to map the two computation phases of SPMMs.To synchronize the granularity of data fetching with the memory access pattern, a novel memory-aware compressed format is proposed for sparse encoding to further reduce the memory transaction and the decoding overhead of ๐ถ 3 ๐๐โ๐.To realize ๐ถ 3 ๐๐โ๐, an
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2fcbfc13-e5f5-4796-823e-470132c0d06fRelated papers
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 ยท 280 citations
- Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationGuowei Zhang, Nithya Attaluri, Joel S. Emer, Daniel SรกnchezASPLOS 2021 ยท 158 citations
- ACES: Accelerating Sparse Matrix Multiplication with Adaptive Execution Flow and Concurrency-Aware Cache OptimizationsXiaoyang Lu, Boyu Long, Xiaoming Chen, Yinhe Han et al.ASPLOS 2024 ยท 13 citations
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 ยท 18 citations
- SPAGHETTI: Streaming Accelerators for Highly Sparse GEMM on FPGAsReza Hojabr, Ali Sedaghati, Amirali Sharifian, Ahmad Khonsari et al.HPCA 2021 ยท 66 citations
