EdgeNN: Efficient Neural Network Inference for CPU-GPU Integrated Edge Devices
Chenyang Zhang, Feng Zhang, Kuangyu Chen, Mingjun Chen, Bingsheng He, Xiaoyong Du
摘要
With the development of the architectures and the growth of AIoT application requirements, data processing on edge has become popular. Neural network inference is widely employed for data analytics on edge devices. This paper extensively explores neural network inference on integrated edge devices and proposes EdgeNN, the first neural network inference solution on CPU-GPU integrated edge devices. EdgeNN has three novel characteristics. First, EdgeNN can adaptively utilize the unified physical memory and conduct the zero-copy optimization. Second, EdgeNN involves a novel inference-targeted inter- and intra-kernel CPU-GPU hybrid execution approach, which co-runs the CPU with the GPU to fully utilize the edge device’s computing resources. Third, EdgeNN adopts a fine-grained adaptive inference tuning approach, which can divide the complicated inference structure into sub-tasks mapped to the CPU and the GPU. Experiments show that on six popular neural network inference tasks, EdgeNN brings an average of 3.97×, 3.12×, and 8.80× speedups to inference on the CPU of the integrated device, inference on a mobile phone CPU, and inference on an edge CPU device. Additionally, it achieves 22.02% time benefits to the direct execution of the original programs. Specifically, 9.93% comes from better utilization of unified memory, and 10.76% comes from CPU-GPU hybrid execution. Besides, EdgeNN can deliver 29.14× and 5.70× higher energy efficiency than the edge CPU and the discrete GPU, respectively. We have made EdgeNN available at https://github.com/ChenyangZhang-cs/EdgeNN.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- HySAE: An Efficient Semantic-Enhanced Representation Learning Model for Knowledge Hypergraph Link PredictionZhao Li, Xin Wang, Jun Zhao, Feng Feng 等WWW 2025 · 被引用 14 次
- Online Container Caching with Late-Warm for IoT Data ProcessingGuopeng Li, Haisheng Tan, Xuan Zhang, Chi Zhang 等ICDE 2024 · 被引用 4 次
相关 Paper
- NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End PerformanceTianyu Jia, Yuhao Ju, Russ Joseph, Jie GuMICRO 2020 · 被引用 21 次
- InfScaler: Enabling Efficient ML Inference Serving on Multi-Accelerator Edge Devices via Asymmetric Auto-ScalingBorui Li, Tiange Xia, Shuai Wang, Shuai WangDAC 2025 · 被引用 2 次
- Graph Neural Networks Automated Design and Deployment on Device-Edge Co-Inference SystemsAo Zhou, Jianlei Yang, Tong Qiao, Yingjie Qi 等DAC 2024 · 被引用 5 次
- Minimizing Latency for Multi-DNN Inference on Resource-Limited CPU-Only Edge DevicesTao Wang, Tuo Shi, Xiulong Liu, Jianping Wang 等INFOCOM 2024 · 被引用 9 次
- AutoScale: Energy Efficiency Optimization for Stochastic Edge Inference Using Reinforcement LearningYoung Geun Kim, Carole-Jean WuMICRO 2020 · 被引用 80 次
