EDA: Energy-Efficient Inter-Layer Model Compilation for Edge DNN Inference Acceleration
Bo Ren Pao, I-Chia Chen, En-Hao Chang, Tsung Tai Yeh
Abstract
Modern handheld devices often employ neural processing units (NPUs) to accelerate deep neural network (DNN) inference applications. Unlike the AI accelerator of a data center, the NPU of an edge device has strict price and energy budgets. In an NPU, DRAM memory consumes much energy due to frequent data movement across on-chip and off-chip memory. The inter-layer operator scheduling has been shown to reduce off-chip memory transactions by reusing DNN operator outputs on the on-chip memory space. However, it often deeply traverses operators and substantially increases off-chip memory data traffic when the on-chip memory of an NPU decreases in size. Consequently, this work creates a DNN model compilation framework called EDA that transparently improves the energy efficiency of an NPU by adjusting the operator traversal depth of the inter-layer operator scheduling in a stacked DNN model. First, EDA transforms a DNN model into a tensor-splitting model. Second, EDA breaks the tensor-splitting model graph into multiple subgraphs. Third, the EDA inter-layer cost model quickly determines the depth of each subgraph. Fourth, EDA properly manages the on-chip shared memory space of an NPU to avoid overusing memory space. Finally, EDA devises the operator grouping method to improve the MAC unit and on-chip memory space utilization. EDA improves the geometric means of and , respectively, in energy efficiency and performance over the NPU with designated memory buffers.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 40cefffb-65c8-4b8a-b979-2e604c3f40faRelated papers
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
- Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise SchedulingYoung H. Oh, Seonghak Kim, Yunho Jin, Sam Son et al.HPCA 2021 · 46 citations
- RESPECT: Reinforcement Learning based Edge Scheduling on Pipelined Coral Edge TPUsJiaqi Yin, Yingjie Li, Daniel Robinson, Cunxi YuDAC 2023 · 9 citations
- Inter-layer Scheduling Space Definition and Exploration for Tiled AcceleratorsJingwei Cai, Yuchen Wei, Zuotong Wu, Sen Peng et al.ISCA 2023 · 67 citations
- Improving Data Reuse in NPU On-chip Memory with Interleaved Gradient Order for DNN TrainingJungwoo Kim, Seonjin Na, Sanghyeon Lee, Sunho Lee et al.MICRO 2023 · 4 citations
