EDA: Energy-Efficient Inter-Layer Model Compilation for Edge DNN Inference Acceleration
Bo Ren Pao, I-Chia Chen, En-Hao Chang, Tsung Tai Yeh
摘要
Modern handheld devices often employ neural processing units (NPUs) to accelerate deep neural network (DNN) inference applications. Unlike the AI accelerator of a data center, the NPU of an edge device has strict price and energy budgets. In an NPU, DRAM memory consumes much energy due to frequent data movement across on-chip and off-chip memory. The inter-layer operator scheduling has been shown to reduce off-chip memory transactions by reusing DNN operator outputs on the on-chip memory space. However, it often deeply traverses operators and substantially increases off-chip memory data traffic when the on-chip memory of an NPU decreases in size. Consequently, this work creates a DNN model compilation framework called EDA that transparently improves the energy efficiency of an NPU by adjusting the operator traversal depth of the inter-layer operator scheduling in a stacked DNN model. First, EDA transforms a DNN model into a tensor-splitting model. Second, EDA breaks the tensor-splitting model graph into multiple subgraphs. Third, the EDA inter-layer cost model quickly determines the depth of each subgraph. Fourth, EDA properly manages the on-chip shared memory space of an NPU to avoid overusing memory space. Finally, EDA devises the operator grouping method to improve the MAC unit and on-chip memory space utilization. EDA improves the geometric means of and , respectively, in energy efficiency and performance over the NPU with designated memory buffers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise SchedulingYoung H. Oh, Seonghak Kim, Yunho Jin, Sam Son 等HPCA 2021 · 被引用 46 次
- RESPECT: Reinforcement Learning based Edge Scheduling on Pipelined Coral Edge TPUsJiaqi Yin, Yingjie Li, Daniel Robinson, Cunxi YuDAC 2023 · 被引用 9 次
- Inter-layer Scheduling Space Definition and Exploration for Tiled AcceleratorsJingwei Cai, Yuchen Wei, Zuotong Wu, Sen Peng 等ISCA 2023 · 被引用 67 次
- Improving Data Reuse in NPU On-chip Memory with Interleaved Gradient Order for DNN TrainingJungwoo Kim, Seonjin Na, Sanghyeon Lee, Sunho Lee 等MICRO 2023 · 被引用 4 次
