Atomic Dataflow based Graph-Level Workload Orchestration for Scalable DNN Accelerators
Shixuan Zheng, Xianjue Zhang, Leibo Liu, Shaojun Wei, Shouyi Yin
摘要
To efficiently deploy state-of-the-art deep neural network (DNN) workloads with growing computational intensity and structural complexity, scalable DNN accelerators have been proposed in recent years, which are featured by multi-tensor engines and distributed on-chip buffers. Such spatial architectures have significantly expanded scheduling space in terms of parallelism and data reuse potentials, which demands for delicate workload orchestration. Previous works on DNN’s hardware mapping problem mainly focus on operator-level loop transformation for single array, which are insufficient for this new challenge. Resource partitioning methods for multi-engines such as CNN-partition and inter-layer pipelining have been studied. However, their intrinsic disadvantages of workload unbalance and pipeline delay still prevent scalable accelerators from releasing full potentials.In this paper, we propose atomic dataflow, a novel graph-level scheduling and mapping approach developed for DNN inference. Instead of partitioning hardware resources into fixed regions and binding each DNN layer to a certain region sequentially, atomic dataflow schedules the DNN computation graph in workload-specific granularity (atoms) to ensure PE-array utilization, supports flexible atom ordering to exploit parallelism, and orchestrates atom-engine mapping to optimize data reuse between spatially connected tensor engines. Firstly, we propose a simulated annealing based atomic tensor generation algorithm to minimize load unbalance. Secondly, we develop a dynamic programming based atomic DAG scheduling algorithm to systematically explore massive ordering potentials. Finally, to facilitate data locality and reduce expensive off-chip memory access, we present mapping and buffering strategies to efficiently utilize distributed on-chip storage. With an automated optimization framework being established, experimental results show significant improvements over baseline approaches in terms of performance, hardware utilization, and energy consumption.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper11
- Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsJingwei Cai, Zuotong Wu, Sen Peng, Yuchen Wei 等HPCA 2024 · 被引用 65 次
- Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators FusionSize Zheng, Siyuan Chen, Peidi Song, Renze Chen 等HPCA 2023 · 被引用 46 次
- TileFlow: A Framework for Modeling Fusion Dataflow via Tree-based AnalysisSize Zheng, Siyuan Chen, Siyuan Gao, Liancheng Jia 等MICRO 2023 · 被引用 31 次
- SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module AcceleratorsMohanad Odema, Luke Chen, Hyoukjun Kwon, Mohammad Abdullah Al FaruqueMICRO 2024 · 被引用 11 次
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsZheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 被引用 9 次
相关 Paper
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- SFD: Towards Segment Fusion Dataflow for Spatial AcceleratorsFuyu Wang, Minghua Shen, Yufei Ding, Nong Xiao 等HPCA 2026
- Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoCSize Zheng, Siyuan Chen, Yun LiangDAC 2023 · 被引用 10 次
- Effectively Scheduling Computational Graphs of Deep Neural Networks toward Their Domain-Specific AcceleratorsJie Zhao, Siyuan Feng, Xiaoqiang Dan, Fei Liu 等OSDI 2023 · 被引用 9 次
- CoSA: Scheduling by Constrained Optimization for Spatial AcceleratorsQijing Huang, Aravind Kalaiah, Minwoo Kang, James Demmel 等ISCA 2021 · 被引用 120 次
