Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoC
Size Zheng, Siyuan Chen, Yun Liang
摘要
The DNN models are now pervasively used for various applications. Meanwhile, the computing hardware has shifted towards heterogeneous system composed of various accelerators. The intertwined complexity of DNN models and hardware makes it challenging for mapping DNN models. Existing mapping frameworks suffer from inefficiencies due to under utilization of computation and bandwidth in heterogeneous SoC. In this paper, we propose COMB, a mapping framework that coordinates the memory and computation and data transfer overhead of heterogeneous accelerators to achieve latency improvement and energy efficiency with two optimizations: dataflow grouping and accelerator mapping. Dataflow grouping maps multiple independent DNN layers to the same accelerator at the same time to spatially share the hardware resources; accelerator mapping finds the optimized placement of the layer groups to accelerators to reduce data transfer overhead. These two optimizations provide a huge design space for heterogeneous DNN mapping. To explore the space efficiently, we present a hybrid scheduling algorithm by combining greedy algorithm and genetic algorithm. In evaluation, COMB achieves 1.28× and 1.37× speedup for latency compared to MAGMA and H2H; COMB also reduces 22.7% and 29.2% energy consumption compared to MAGMA and H2H.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MAGIS: Memory Optimization via Coordinated Graph Transformation and Scheduling for DNNRenze Chen, Zijian Ding, Size Zheng, Chengrui Zhang 等ASPLOS 2024 · 被引用 14 次
- Ev-Edge: Efficient Execution of Event-based Vision Algorithms on Commodity Edge PlatformsShrihari Sridharan, Surya Selvam, Kaushik Roy, Anand RaghunathanDAC 2024 · 被引用 3 次
- AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIMYuanpeng Zhang, Xing Hu, Xi Chen, Zhihang Yuan 等ISCA 2025 · 被引用 1 次
它引用的顶会 Paper10
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen 等ASPLOS 2020 · 被引用 171 次
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 被引用 150 次
- Heterogeneous Dataflow Accelerators for Multi-DNN WorkloadsHyoukjun Kwon, Liangzhen Lai, Michael Pellauer, Tushar Krishna 等HPCA 2021 · 被引用 143 次
- A Multi-Neural Network Acceleration ArchitectureEunjin Baek, Dongup Kwon, Jangwoo KimISCA 2020 · 被引用 110 次
- TENET: A Framework for Modeling Tensor Dataflow Based on Relation-centric NotationLiqiang Lu, Naiqing Guan, Yuyue Wang, Liancheng Jia 等ISCA 2021 · 被引用 82 次
相关 Paper
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsGuan Shen, Jieru Zhao, Zeke Wang, Zhe Lin 等DAC 2023 · 被引用 5 次
- Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-ChipsIsmet Dagli, Mehmet E. BelviranliPPoPP 2024 · 被引用 18 次
- H2H: heterogeneous model to heterogeneous system mapping with computation and communication awarenessXinyi Zhang, Cong Hao, Peipei Zhou, Alex K. Jones 等DAC 2022 · 被引用 21 次
- CODO: An Automated Compiler for Comprehensive Dataflow OptimizationWeichuang Zhang, Yiquan Wang, Xinzhou Zhang, Chi Zhang 等ISCA 2026
- Map-and-Conquer: Energy-Efficient Mapping of Dynamic Neural Nets onto Heterogeneous MPSoCsHalima Bouzidi, Mohanad Odema, Hamza Ouarnoughi, Smaïl Niar 等DAC 2023 · 被引用 14 次
