SegFold: Accelerating Sparse Gemm with a Fine-Grained Dynamic Dataflow
Xinrui Wu, Hanyu Wang, Jason Cong, Tony Nowatzki
Abstract
Generalized sparse matrix-matrix multiplication (SpGEMM) is critical in many domains. Existing CPUs and GPUs, as well as specialized accelerators, rely on static dataflows (e.g., inner product, outer product, Gustavson, etc.). Each static dataflow sacrifices some data reuse opportunities and imposes constraints on load balance. To address this inefficiency, we extend the typical SpGEMM dataflows by considering dynamism. Specifically, we add finegrained dynamic scheduling to optimize reuse and reduce resource contention. We also develop dynamic remapping of partially completed work to improve load balance and parallelism. These ideas are formalized into a specific dataflow called Segment. To demonstrate Segment, we codesign a SpGEMM accelerator called SegFold. SegFold includes a memory controller that identifies fine-grained reuse opportunities in a local window of the stationary input array and exploits them through dynamic work assignment. It also incorporates a merge network that routes inputs to appropriate processing elements (PEs) for computation while dynamically remapping the work assigned to each PE to balance load. Across diverse densities and matrix sizes, SegFold achieves a geometric-mean 1.95× speedup over state-of-the-art SpGEMM accelerators and over the best static dataflow configuration, demonstrating that adding dynamism to the dataflow design space unlocks reuse and load-balance gains that no static scheduling choice can achieve in isolation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f9b578f-f9b4-4c75-9452-c1b0b25d31beBuilds on19
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload RebalancingTong Geng, Ang Li, Runbin Shi, Chunshu Wu et al.MICRO 2020 · 299 citations
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
- MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise ProductNitish Kumar Srivastava, Hanchen Jin, Jie Liu, David H. Albonesi et al.MICRO 2020 · 223 citations
- Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationGuowei Zhang, Nithya Attaluri, Joel S. Emer, Daniel SánchezASPLOS 2021 · 158 citations
Related papers
- Spada: Accelerating Sparse Matrix Multiplication with Adaptive DataflowZhiyao Li, Jiaxiang Li, Taijie Chen, Dimin Niu et al.ASPLOS 2023 · 59 citations
- SLAWS: Spatial Locality Analysis and Workload Orchestration for Sparse Matrix MultiplicationGuoyu Li, Zheng Guan, Beichen Zhang, Jun Yu et al.ASPLOS 2026
- SPAGHETTI: Streaming Accelerators for Highly Sparse GEMM on FPGAsReza Hojabr, Ali Sedaghati, Amirali Sharifian, Ahmad Khonsari et al.HPCA 2021 · 66 citations
- SPADE: A Flexible and Scalable Accelerator for SpMM and SDDMMGerasimos Gerogiannis, Serif Yesil, Damitha Lenadora, Dingyuan Cao et al.ISCA 2023 · 27 citations
- SpaHet: A Software/Hardware Co-design for Accelerating Heterogeneous-Sparsity based Sparse Matrix MultiplicationHaoqin Huang, Pengcheng Yao, Zhaozeng An, Yufei Sun et al.DAC 2024 · 3 citations
