Vectorizing Sparse Matrix Computations with Partially-Strided Codelets
Kazem Cheshmi, Zachary Cetinic, Maryam Mehri Dehnavi
摘要
The compact data structures and irregular computation patterns in sparse matrix computations introduce challenges to vectorizing these codes. Available approaches primarily vectorize strided regions of computation in a sparse code. They also reorganize data and computations, at a cost, to increase the number of strided regions. In this work, we propose a locality-based codelet mining (LCM) algorithm that efficiently searches for strided and partially strided regions in sparse matrix computations for vectorization. We also present a classification of partially strided codelets along with a differentiation-based approach to generate codelets from memory accesses in the sparse computation. LCM is implemented as an inspector-executor framework called LCM I/E. It generates vectorized code for the sparse matrix-vector multiplication (SpMV) and sparse matrix times dense matrix (SpMM), kernels with parallel outermost loops, and kernels with loop-carried dependence, specifically the sparse triangular solver (SpTRSV). We demonstrate the performance of the LCM I/E-generated code for SpMV/SpMM on a set of 789 real matrices (0.1-330M nonzeros) and SpTRSV on a set of 132 symmetric positive definite matrices. LCM I/E outperforms the highly specialized library MKL with an average speedup of 1.67×, 4.1×, 1.75× for SpMV, SpTRSV, and SpMM, respectively. For the same matrices, LCM I/E outperforms the state-of-the-art inspector-executor framework Sympiler [1] for the SpTRSV kernel with an average speedup of 1.9×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Register Tiling for Unstructured Sparsity in Neural Network InferenceLucas Wilkinson, Kazem Cheshmi, Maryam Mehri DehnaviPLDI 2023 · 被引用 17 次
- Runtime Composition of Iterations for Fusing Loop-carried Sparse DependenceKazem Cheshmi, Michelle Strout, Maryam Mehri DehnaviSC 2023 · 被引用 9 次
- Mille-feuille: A Tile-Grained Mixed Precision Single-Kernel Conjugate Gradient Solver on GPUsDechuang Yang, Yuxuan Zhao, Yiduo Niu, Weile Jia 等SC 2024 · 被引用 8 次
- Modular Construction and Optimization of the UZP Sparse Format for SpMV on CPUsAlonso Rodríguez-Iglesias, Santoshkumar T. Tongli, Emily Tucker, Louis-Noël Pouchet 等PLDI 2025 · 被引用 1 次
- GALA: A High Performance Graph Neural Network Acceleration LAnguage and CompilerDamitha Lenadora, Nikhil Jayakumar, Chamika Sudusinghe, Charith MendisOOPSLA 2025 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- SpV8: Pursuing Optimal Vectorization and Regular Computation Pattern in SpMVChenyang Li, Tian Xia, Wenzhe Zhao, Nanning Zheng 等DAC 2021 · 被引用 17 次
- WISE: Predicting the Performance of Sparse Matrix Vector Multiplication with Machine LearningSerif Yesil, Azin Heidarshenas, Adam Morrison, Josep TorrellasPPoPP 2023 · 被引用 33 次
- Efficiently running SpMV on long vector architecturesConstantino Gómez, Filippo Mantovani, Erich Focht, Marc CasasPPoPP 2021 · 被引用 48 次
- PANA: A Fine-Grained Runtime-Adaptive Load Balancing for Parallel SpMV on Multicore CPUsHaodong Bian, Youhui Zhang, Xiang Fei, Jianqiang Huang 等PPoPP 2026 · 被引用 2 次
- A Hardware-Software Design Framework for SpMV Acceleration with Flexible Access Pattern PortfolioZhenyu Wu, Maolin Wang, Hayden Kwok-Hay SoHPCA 2025 · 被引用 1 次
