Azul: An Accelerator for Sparse Iterative Solvers Leveraging Distributed On-Chip Memory
Axel Feldmann, Courtney Golden, Yifan Yang, Joel S. Emer, Daniel Sánchez
Abstract
Solving sparse systems of linear equations is a fundamental primitive in many numeric algorithms. Iterative solvers provide an efficient way of solving large, highly sparse systems. However, iterative solvers are inefficient on existing architectures because they perform computations with (1) poor short-term reuse, which causes frequent off-chip memory traffic; and (2) challenaing data dependences, which limit parallelism. We present Azul, a hardware accelerator that achieves high arithmetic intensity by keeping data in distributed on-chip SRAM. Azul is organized as a grid of tiles, each with a small memory and a simple processing element (PE). This enables keeping solver data on-chip across iterations, achieving high reuse. We present a novel scheduling algorithm that maps data and computation across PEs to avoid communication bottlenecks while achieving high parallelism, and a specialized PE that achieves high utilization of arithmetic units. When tested on a representative set of matrices for sparse iterative solvers, Azul is gmean 217 × faster than state-of-the art GPU implementations, 159× faster than a previously proposed accelerator for sparse iterative solvers, and 90 × faster than a previously proposed distributed-SRAM accelerator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 019373fd-98df-4fc4-94df-30d2a0fc21cbBuilds on5
- ALRESCHA: A Lightweight Reconfigurable Sparse-Computation AcceleratorBahar Asgari, Ramyad Hadidi, Tushar Krishna, Hyesoon Kim et al.HPCA 2020 · 63 citations
- Dalorex: A Data-Local Program Execution and Architecture for Memory-bound ApplicationsMarcelo Orenes-Vera, Esin Tureci, David Wentzlaff, Margaret MartonosiHPCA 2023 · 23 citations
- FDMAX: An Elastic Accelerator Architecture for Solving Partial Differential EquationsJiajun Li, Yuxuan Zhang, Hao Zheng, Ke WangISCA 2023 · 16 citations
- RSQP: Problem-specific Architectural Customization for Accelerated Convex Quadratic OptimizationMaolin Wang, Ian McInerney, Bartolomeo Stellato, Stephen P. Boyd et al.ISCA 2023 · 9 citations
- Spatula: A Hardware Accelerator for Sparse Matrix FactorizationAxel Feldmann, Daniel SánchezMICRO 2023 · 9 citations
Related papers
- Quartz: A Reconfigurable, Distributed-Memory Accelerator for Sparse ApplicationsCourtney Golden, Axel Feldmann, Joel S. Emer, Daniel SánchezMICRO 2025 · 1 citation
- HYTE: Flexible Tiling for Sparse Accelerators via Hybrid Static-Dynamic ApproachesXintong Li, Zhiyao Li, Mingyu GaoISCA 2025 · 2 citations
- HiSpTRSV: Exploring Tile-Level Parallelism for SpTRSV Acceleration on FPGAsFan Sun, Fang Dong, Dian ShenDAC 2025
- HotTiles: Accelerating SpMM with Heterogeneous Accelerator ArchitecturesGerasimos Gerogiannis, Sriram Aananthakrishnan, Josep Torrellas, Ibrahim HurHPCA 2024 · 19 citations
- AmgR: Algebraic Multigrid Accelerated on ReRAMMingjia Fan, Xiaotian Tian, Yintao He, Junxian Li et al.DAC 2023 · 8 citations
