SARA: Scaling a Reconfigurable Dataflow Accelerator
Yaqi Zhang, Nathan Zhang, Tian Zhao, Matt Vilim, Muhammad Shahbaz, Kunle Olukotun
Abstract
The need for speed in modern data-intensive work-loads and the rise of "dark silicon" in the semiconductor industry are pushing for larger, faster, and more energy and area-efficient architectures, such as Reconfigurable Dataflow Accelerators (RDAs). Nevertheless, challenges remain in developing mechanisms to effectively utilize the compute power of these large-scale RDAs. To address these challenges, we present SARA, a compiler that employs a novel mapping strategy to efficiently utilize large-scale RDAs. Starting from a single-threaded imperative abstraction, SARA spatially maps a program onto RDA's distributed resources, exploiting dataflow parallelism within and across hyperblocks to saturate the compute throughput of an RDA. SARA introduces (a) compiler-managed memory consistency (CMMC), a control paradigm that hierarchically pipelines a nested and data-dependent control-flow graph onto a dataflow architecture, and (b) a compilation flow that decomposes the program graph across distributed heterogeneous resources to hide low-level RDA constraints from programmers. Our evaluation shows that SARA achieves close to perfect performance scaling on a recently proposed RDA—Plasticine. Over a mix of deep-learning, graph-processing, and streaming applications, SARA achieves a 1.9× geo-mean speedup over a Tesla V100 GPU using only 12% of the silicon area.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 54d53d63-3785-47a2-ab41-39fff7ad3f9dCited by top-tier papers14
- Taurus: a data plane architecture for per-packet MLTushar Swamy, Alexander Rucker, Muhammad Shahbaz, Ishan Gaur et al.ASPLOS 2022 · 94 citations
- AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstractionSize Zheng, Renze Chen, Anjiang Wei, Yicheng Jin et al.ISCA 2022 · 63 citations
- Capstan: A Vector RDA for SparsityAlexander Rucker, Matthew Vilim, Tian Zhao, Yaqi Zhang et al.MICRO 2021 · 37 citations
- The Sparse Abstract MachineOlivia Hsu, Maxwell Strange, Ritvik Sharma, Jaeyeon Won et al.ASPLOS 2023 · 37 citations
- TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based AnalysisSize Zheng, Siyuan Chen, Siyuan Gao, Liancheng Jia et al.MICRO 2023 · 31 citations
Related papers
- Revet: A Language and Compiler for Dataflow ThreadsAlexander C. Rucker, Shiv Sundram, Coleman Smith, Matthew Vilim et al.HPCA 2024 · 3 citations
- LISA: Graph Neural Network based Portable Mapping on Spatial AcceleratorsZhaoying Li, Dan Wu, Dhananjaya Wijerathne, Tulika MitraHPCA 2022 · 43 citations
- PANORAMA: divide-and-conquer approach for mapping complex loop kernels on CGRADhananjaya Wijerathne, Zhaoying Li, Thilini Kaushalya Bandara, Tulika MitraDAC 2022 · 18 citations
- FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming DataflowRubens Lacouture, Nathan Zhang, Ritvik Sharma, Marco Siracusa et al.ASPLOS 2026 · 1 citation
- MapZero: Mapping for Coarse-grained Reconfigurable Architectures with Reinforcement Learning and Monte-Carlo Tree SearchXiangyu Kong, Yi Huang, Jianfeng Zhu, Xingchen Man et al.ISCA 2023 · 30 citations
