Quartz: A Reconfigurable, Distributed-Memory Accelerator for Sparse Applications
Courtney Golden, Axel Feldmann, Joel S. Emer, Daniel Sánchez
Abstract
Iterative sparse matrix computations lie at the heart of many scientific computing and graph analytics algorithms.On conventional systems, their irregular memory accesses and low arithmetic intensity create challenging memory bandwidth bottlenecks.To overcome such bottlenecks, distributed-SRAM architectures are structured as an array of tiles, each with a processing element (PE) and a small local memory, to achieve very high aggregate memory bandwidth.However, current distributed-SRAM architectures suffer from either poor programmability due to over-specialized PEs or poor compute performance due to inefficient general-purpose PEs.We propose Quartz, a new architecture that uses short dataflow tasks and reconfigurable PEs in a distributed-SRAM system to deliver both high performance and high programmability.Unlike traditional sparse CGRAs or on-die reconfigurable engines, Quartz allows reconfigurable compute to be highly utilized and scaled by (1) providing high memory bandwidth to each processing element and (2) introducing a task-level dataflow execution model that fits this new setting.Our execution model dynamically reconfigures each tile's PE in response to inter-tile messages to execute tasks on local data.This execution model enables fine-grained data partitioning across tiles.To make execution efficient, we explore novel data partitioning techniques that use graph and hypergraph partitioning to minimize network traffic and balance load in the face of both static-static and static-dynamic operand sparsity.To ensure programmability, we show how a wide range of Einsum-expressible computations and flexible data distributions can be systematically captured in small tasks for execution on Quartz.Quartz's architecture, data partitioning techniques, and programming model together achieve gmean 21.4× speedup over a prior state-of-the-art system for six different iterative sparse applications from scientific computing and graph analytics.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ae904259-5939-43f8-8edd-ce293b43773dRelated papers
- Azul: An Accelerator for Sparse Iterative Solvers Leveraging Distributed On-Chip MemoryAxel Feldmann, Courtney Golden, Yifan Yang, Joel S. Emer et al.MICRO 2024 · 7 citations
- SARA: Scaling a Reconfigurable Dataflow AcceleratorYaqi Zhang, Nathan Zhang, Tian Zhao, Matt Vilim et al.ISCA 2021 · 58 citations
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 23 citations
- Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix MultiplicationIsuru Ranawaka, Md Taufique Hussain, Charles Block, Gerasimos Gerogiannis et al.SC 2024 · 5 citations
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim et al.HPCA 2025 · 5 citations
