Spatula: A Hardware Accelerator for Sparse Matrix Factorization
Axel Feldmann, Daniel Sánchez
摘要
Solving sparse systems of linear equations is a crucial component in many science and engineering problems, like simulating physical systems. Sparse matrix factorization dominates a large class of these solvers. Efficient factorization algorithms have two key properties that make them challenging for existing architectures: they consist of small tasks that are structured and compute-intensive, and sparsity induces long chains of data dependences among these tasks. Data dependences make GPUs struggle, while CPUs and prior sparse linear algebra accelerators also suffer from low compute throughput.
We present Spatula, an architecture for accelerating sparse matrix factorization algorithms. Spatula hardware combines systolic processing elements that execute structured tasks at high throughput with a flexible scheduler that handles challenging data dependences. Spatula enables a novel scheduling algorithm that avoids stalls and load imbalance while reducing data movement, achieving high compute utilization. As a result, Spatula outperforms a GPU running the state-of-the-art sparse Cholesky and LU factorization implementations by gmean 47× across a wide range of matrices, and by up to thousands of times on some challenging matrices.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache QuantizationMinsu Kim, Seongmin Hong, Ryeowook Ko, Soongyu Choi 等ISCA 2025 · 被引用 17 次
- Azul: An Accelerator for Sparse Iterative Solvers Leveraging Distributed On-Chip MemoryAxel Feldmann, Courtney Golden, Yifan Yang, Joel S. Emer 等MICRO 2024 · 被引用 7 次
- SuperNoVA: Algorithm-Hardware Co-Design for Resource-Aware SLAMSeah Kim, Roger Hsiao, Borivoje Nikolic, James Demmel 等ASPLOS 2025 · 被引用 7 次
- Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsQizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng 等HPCA 2025 · 被引用 2 次
- Multi-Issue Butterfly Architecture for Sparse Convex Quadratic ProgrammingMaolin Wang, Ian McInerney, Bartolomeo Stellato, Fengbin Tu 等MICRO 2024 · 被引用 2 次
它引用的顶会 Paper14
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 被引用 280 次
- Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationGuowei Zhang, Nithya Attaluri, Joel S. Emer, Daniel SánchezASPLOS 2021 · 被引用 158 次
- Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor ComputationsNitish Kumar Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong 等HPCA 2020 · 被引用 121 次
- SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorXinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak 等HPCA 2021 · 被引用 111 次
- Dual-side Sparse Tensor CoreYang Wang, Chen Zhang, Zhiqiang Xie, Cong Guo 等ISCA 2021 · 被引用 109 次
相关 Paper
- SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUsJianqi Zhao, Yao Wen, Yuchen Luo, Zhou Jin 等DAC 2021 · 被引用 30 次
- ALRESCHA: A Lightweight Reconfigurable Sparse-Computation AcceleratorBahar Asgari, Ramyad Hadidi, Tushar Krishna, Hyesoon Kim 等HPCA 2020 · 被引用 63 次
- Addressing Irregular Patterns of Matrix Computations on GPUs and Their Impact on Applications Powered by Sparse Direct SolversAhmad Abdelfattah, Pieter Ghysels, Wajih Boukaram, Stanimire Tomov 等SC 2022 · 被引用 4 次
- Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained SchedulingJie Ren, Tingxuan Zhong, Yuxi Hong, Guofeng Feng 等SC 2025 · 被引用 1 次
- SPAGHETTI: Streaming Accelerators for Highly Sparse GEMM on FPGAsReza Hojabr, Ali Sedaghati, Amirali Sharifian, Ahmad Khonsari 等HPCA 2021 · 被引用 66 次
