Runtime Composition of Iterations for Fusing Loop-carried Sparse Dependence
Kazem Cheshmi, Michelle Strout, Maryam Mehri Dehnavi
摘要
Dependence between iterations in sparse computations causes inefficient use of memory and computation resources. This paper proposes sparse fusion, a technique that generates efficient parallel code for the combination of two sparse matrix kernels, where at least one of the kernels has loop-carried dependencies. Existing implementations optimize individual sparse kernels separately. However, this approach leads to synchronization overheads and load imbalance due to the irregular dependence patterns of sparse kernels, as well as inefficient cache usage due to their irregular memory access patterns. Sparse fusion uses a novel inspection strategy and code transformation to generate parallel fused code optimized for data locality and load balance. Sparse fusion outperforms the best of unfused implementations using ParSy and MKL by an average of 4.2× and is faster than the best of fused implementations using existing scheduling algorithms, such as LBC, DAGP, and wavefront by an average of 4× for various kernel combinations.
• Software and its engineering → Runtime environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SPLAT: A Framework for Optimised GPU Code-Generation for SParse reguLar ATtentionAhan Gupta, Yueming Yuan, Devansh Jain, Yuhao Ge 等OOPSLA 2025 · 被引用 3 次
- Adaptive Algebraic Reuse of Reordering in Cholesky Factorizations with Dynamic Sparsity PatternsBehrooz Zarebavani, Danny M. Kaufman, David I. W. Levin, Maryam Mehri DehnaviSIGGRAPH 2025 · 被引用 1 次
- Modular Construction and Optimization of the UZP Sparse Format for SpMV on CPUsAlonso Rodríguez-Iglesias, Santoshkumar T. Tongli, Emily Tucker, Louis-Noël Pouchet 等PLDI 2025 · 被引用 1 次
- GALA: A High Performance Graph Neural Network Acceleration LAnguage and CompilerDamitha Lenadora, Nikhil Jayakumar, Chamika Sudusinghe, Charith MendisOOPSLA 2025 · 被引用 1 次
它引用的顶会 Paper5
- Lessons Learned from the Chameleon TestbedKate Keahey, Jason Anderson, Zhuo Zhen, Pierre Riteau 等USENIX ATC 2020 · 被引用 398 次
- SparseTIR: Composable Abstractions for Sparse Compilation in Deep LearningZihao Ye, Ruihang Lai, Junru Shao, Tianqi Chen 等ASPLOS 2023 · 被引用 86 次
- NASOQ: numerically accurate sparsity-oriented QP solverKazem Cheshmi, Danny M. Kaufman, Shoaib Kamil, Maryam Mehri DehnaviSIGGRAPH 2020 · 被引用 29 次
- Register Tiling for Unstructured Sparsity in Neural Network InferenceLucas Wilkinson, Kazem Cheshmi, Maryam Mehri DehnaviPLDI 2023 · 被引用 17 次
- Vectorizing Sparse Matrix Computations with Partially-Strided CodeletsKazem Cheshmi, Zachary Cetinic, Maryam Mehri DehnaviSC 2022 · 被引用 4 次
相关 Paper
- Path-sensitive sparse analysis without path conditionsQingkai Shi, Peisen Yao, Rongxin Wu, Charles ZhangPLDI 2021 · 被引用 24 次
- Lightweight and Locality-Aware Composition of Black-Box SubroutinesManya Bansal, Dillon Sharlet, Jonathan Ragan-Kelley, Saman P. AmarasinghePLDI 2025
- SpaceFusion: Advanced Deep Learning Operator Fusion via Space-Mapping GraphLiang Zhu, Jianguo Yao, Haibing GuanEuroSys 2025 · 被引用 3 次
- Compilation of Shape Operators on Sparse ArraysAlexander J. Root, Bobby Yan, Peiming Liu, Christophe Gyurgyik 等OOPSLA 2024 · 被引用 3 次
- SpV8: Pursuing Optimal Vectorization and Regular Computation Pattern in SpMVChenyang Li, Tian Xia, Wenzhe Zhao, Nanning Zheng 等DAC 2021 · 被引用 17 次
