Azul: An Accelerator for Sparse Iterative Solvers Leveraging Distributed On-Chip Memory
Axel Feldmann, Courtney Golden, Yifan Yang, Joel S. Emer, Daniel Sánchez
摘要
Solving sparse systems of linear equations is a fundamental primitive in many numeric algorithms. Iterative solvers provide an efficient way of solving large, highly sparse systems. However, iterative solvers are inefficient on existing architectures because they perform computations with (1) poor short-term reuse, which causes frequent off-chip memory traffic; and (2) challenaing data dependences, which limit parallelism. We present Azul, a hardware accelerator that achieves high arithmetic intensity by keeping data in distributed on-chip SRAM. Azul is organized as a grid of tiles, each with a small memory and a simple processing element (PE). This enables keeping solver data on-chip across iterations, achieving high reuse. We present a novel scheduling algorithm that maps data and computation across PEs to avoid communication bottlenecks while achieving high parallelism, and a specialized PE that achieves high utilization of arithmetic units. When tested on a representative set of matrices for sparse iterative solvers, Azul is gmean 217 × faster than state-of-the art GPU implementations, 159× faster than a previously proposed accelerator for sparse iterative solvers, and 90 × faster than a previously proposed distributed-SRAM accelerator.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- ALRESCHA: A Lightweight Reconfigurable Sparse-Computation AcceleratorBahar Asgari, Ramyad Hadidi, Tushar Krishna, Hyesoon Kim 等HPCA 2020 · 被引用 63 次
- Dalorex: A Data-Local Program Execution and Architecture for Memory-bound ApplicationsMarcelo Orenes-Vera, Esin Tureci, David Wentzlaff, Margaret MartonosiHPCA 2023 · 被引用 23 次
- FDMAX: An Elastic Accelerator Architecture for Solving Partial Differential EquationsJiajun Li, Yuxuan Zhang, Hao Zheng, Ke WangISCA 2023 · 被引用 16 次
- RSQP: Problem-specific Architectural Customization for Accelerated Convex Quadratic OptimizationMaolin Wang, Ian McInerney, Bartolomeo Stellato, Stephen P. Boyd 等ISCA 2023 · 被引用 9 次
- Spatula: A Hardware Accelerator for Sparse Matrix FactorizationAxel Feldmann, Daniel SánchezMICRO 2023 · 被引用 9 次
相关 Paper
- Quartz: A Reconfigurable, Distributed-Memory Accelerator for Sparse ApplicationsCourtney Golden, Axel Feldmann, Joel S. Emer, Daniel SánchezMICRO 2025 · 被引用 1 次
- HYTE: Flexible Tiling for Sparse Accelerators via Hybrid Static-Dynamic ApproachesXintong Li, Zhiyao Li, Mingyu GaoISCA 2025 · 被引用 2 次
- HiSpTRSV: Exploring Tile-Level Parallelism for SpTRSV Acceleration on FPGAsFan Sun, Fang Dong, Dian ShenDAC 2025
- HotTiles: Accelerating SpMM with Heterogeneous Accelerator ArchitecturesGerasimos Gerogiannis, Sriram Aananthakrishnan, Josep Torrellas, Ibrahim HurHPCA 2024 · 被引用 19 次
- AmgR: Algebraic Multigrid Accelerated on ReRAMMingjia Fan, Xiaotian Tian, Yintao He, Junxian Li 等DAC 2023 · 被引用 8 次
