AxoNN: energy-aware execution of neural network inference on multi-accelerator heterogeneous SoCs
Ismet Dagli, Alexander Cieslewicz, Jedidiah McClurg, Mehmet E. Belviranli
摘要
The energy and latency demands of critical workload execution, such as object detection, in embedded systems vary based on the physical system state and other external factors. Many recent mobile and autonomous System-on-Chips (SoC) embed a diverse range of accelerators with unique power and performance characteristics. The execution flow of the critical workloads can be adjusted to span into multiple accelerators so that the trade-off between performance and energy fits to the dynamically changing physical factors.
In this study, we propose running neural network (NN) inference on multiple accelerators of an SoC. Our goal is to enable an energy-performance trade-off with an by distributing layers in a NN between a performance-and a power-efficient accelerator. We first provide an empirical modeling methodology to characterize execution and inter-layer transition times. We then find an optimal layers-to-accelerator mapping by representing the trade-off as a linear programming optimization constraint. We evaluate our approach on the NVIDIA Xavier AGX SoC with commonly used NN models. We use the Z3 SMT solver to find schedules for different energy consumption targets, with up to 98% prediction accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-ChipsIsmet Dagli, Mehmet E. BelviranliPPoPP 2024 · 被引用 18 次
- Map-and-Conquer: Energy-Efficient Mapping of Dynamic Neural Nets onto Heterogeneous MPSoCsHalima Bouzidi, Mohanad Odema, Hamza Ouarnoughi, Smaïl Niar 等DAC 2023 · 被引用 14 次
- HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML PlatformsJosse Van Delm, Maarten Vandersteegen, Alessio Burrello, Giuseppe Maria Sarda 等DAC 2023 · 被引用 10 次
它引用的顶会 Paper3
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee 等OSDI 2020 · 被引用 286 次
- HetPipe: Enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data ParallelismJay H. Park, Gyeongchan Yun, Chang M. Yi, Nguyen T. Nguyen 等USENIX ATC 2020 · 被引用 178 次
- PCCS: Processor-Centric Contention-aware Slowdown Model for Heterogeneous System-on-ChipsYuanchao Xu, Mehmet Esat Belviranli, Xipeng Shen, Jeffrey S. VetterMICRO 2021 · 被引用 15 次
相关 Paper
- NeuOS: A Latency-Predictable Multi-Dimensional Optimization Framework for DNN-driven Autonomous SystemsSoroush Bateni, Cong LiuUSENIX ATC 2020 · 被引用 49 次
- Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoCSize Zheng, Siyuan Chen, Yun LiangDAC 2023 · 被引用 10 次
- ALERT: Accurate Learning for Energy and TimelinessChengcheng Wan, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann 等USENIX ATC 2020 · 被引用 15 次
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsGuan Shen, Jieru Zhao, Zeke Wang, Zhe Lin 等DAC 2023 · 被引用 5 次
- Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator ExecutionTanmoy Sen, Haiying Shen, Anand Padmanabha IyerEuroSys 2025 · 被引用 2 次
