The Dataflow Abstract Machine Simulator Framework
Nathan Zhang, Rubens Lacouture, Gina Sohn, Paul Mure, Qizheng Zhang, Fredrik Kjolstad, Kunle Olukotun
摘要
The growing interest in novel dataflow architectures and streaming execution paradigms has created the need for a simulator optimized for modeling dataflow systems.
To fill this need, we present three new techniques that make it feasible to simulate complex systems consisting of thousands of components. First, we introduce an interface based on Communicating Sequential Processes which allows users to simultaneously describe functional and timing characteristics. Second, we introduce a scalable point-to-point synchronization scheme that avoids global synchronization. Finally, we demonstrate a technique to exploit slack in the simulated system, such as FIFOs, to increase simulation parallelism.
We implement these techniques in the Dataflow Abstract Machine (DAM), a parallel simulator framework for dataflow systems. We demonstrate the benefits of using DAM by highlighting three case studies using the framework. First, we use DAM directly as an exploration tool for streaming algorithms on dataflow hardware. We simulate two different implementations of the attention algorithm used in large language models, and use DAM to show that the second implementation only requires a constant amount of local memory. Second, we re-implement a simulator for a sparse tensor algebra accelerator, resulting in 57% less code and a simulation speedup of up to four orders of magnitude. Finally, we demonstrate a general technique for timemultiplexing real hardware to simulate multiple virtual copies of the hardware using DAM.
Modern applications such as large language models (LLMs), data analytics, and sparse machine learning have ignited a flurry of research in both dataflow architectures, such as Reconfigurable Dataflow Accelerators (RDAs) and Coarse-Grained Reconfigurable Arrays (CGRAs) [9], [16], [26], [42], [43], [45], [47], and streaming abstractions such as the Sparse Abstract Machine [29]. To explore the functional behavior and performance characteristics of their proposed dataflow systems, many researchers develop bespoke simulators.
To simulate dataflow systems, a framework must (1) support communication among thousands of coupled units, (2) model fine-grained channel behaviors, and (3) simulate heterogenous user-defined models. Furthermore, the scale of these systems demands efficient parallelization. Providing these properties in sequential simulation is straightforward; however, no efficient method exists for parallel simulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationJian Weng, Boyang Han, Derui Gao, Ruijie Gao 等ISCA 2025 · 被引用 1 次
- Streaming Tensor Programs: A Streaming Abstraction for Dynamic ParallelismGina Sohn, Genghan Zhang, Konstantin Hoßfeld, Jungwoo Kim 等ASPLOS 2026 · 被引用 1 次
- FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming DataflowRubens Lacouture, Nathan Zhang, Ritvik Sharma, Marco Siracusa 等ASPLOS 2026 · 被引用 1 次
- Adaptive Self-improvement LLM Agentic System for ML Library DevelopmentGenghan Zhang, Weixin Liang, Olivia Hsu, Kunle OlukotunICML 2025
- Fast End-to-End Performance Simulation of Accelerated Hardware-Software StacksJiacheng Ma, Jonas Kaufmann, Emilien Guandalino, Rishabh R. Iyer 等SOSP 2025
它引用的顶会 Paper5
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Fifer: Practical Acceleration of Irregular Applications on Reconfigurable ArchitecturesQuan M. Nguyen, Daniel SánchezMICRO 2021 · 被引用 60 次
- The Sparse Abstract MachineOlivia Hsu, Maxwell Strange, Ritvik Sharma, Jaeyeon Won 等ASPLOS 2023 · 被引用 37 次
- BaCO: A Fast and Portable Bayesian Compiler Optimization FrameworkErik Orm Hellsten, Artur L. F. Souza, Johannes Lenfers, Rubens Lacouture 等ASPLOS 2023 · 被引用 22 次
相关 Paper
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMsHanchen Ye, Deming ChenMICRO 2025 · 被引用 5 次
- A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC WorkloadsNicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni 等SC 2025 · 被引用 1 次
- Capstan: A Vector RDA for SparsityAlexander Rucker, Matthew Vilim, Tian Zhao, Yaqi Zhang 等MICRO 2021 · 被引用 37 次
- Sigma: Compiling Einstein Summations to Locality-Aware DataflowTian Zhao, Alexander Rucker, Kunle OlukotunASPLOS 2023 · 被引用 3 次
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 被引用 23 次
