Distributed Inference with Deep Learning Models across Heterogeneous Edge Devices
Chenghao Hu, Baochun Li
摘要
Recent years witnessed an increasing research attention in deploying deep learning models on edge devices for inference. Due to limited capabilities and power constraints, it may be necessary to distribute the inference workload across multiple devices. Existing mechanisms divided the model across edge devices with the assumption that deep learning models are constructed with a chain of layers. In reality, however, modern deep learning models are more complex, involving a directed acyclic graph (DAG) rather than a chain of layers.
In this paper, we present EdgeFlow, a new distributed inference mechanism designed for general DAG structured deep learning models. Specifically, EdgeFlow partitions model layers into independent execution units with a new progressive model partitioning algorithm. By producing near-optimal model partitions, our new algorithm seeks to improve the run-time performance of distributed inference as these partitions are distributed across the edge devices. During inference, EdgeFlow orchestrates the intermediate results flowing through these units to fulfill the complicated layer dependencies. We have implemented Edge-Flow based on PyTorch, and evaluated it with state-of-theart deep learning models in different structures. The results show that EdgeFlow reducing the inference latency by up to 40.2% compared with other approaches, which demonstrates the effectiveness of our design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AIZhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang 等INFOCOM 2025 · 被引用 3 次
- ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer InferenceXiao Liu, Lijun Zhang, Deepak Ganesan, Hui GuanICML 2026
- Hierarchical Inference with an Offload QueueSrinivas Nomula, Parimal Parag, Ayalvadi Ganesh, Jaya Prakash ChampatiINFOCOM 2026
相关 Paper
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo 等UbiComp 2020 · 被引用 77 次
- Resource-aware Deployment of Dynamic DNNs over Multi-tiered Interconnected SystemsChetna Singhal, Yashuo Wu, Francesco Malandrino, Marco Levorato 等INFOCOM 2024 · 被引用 15 次
- AGO: Boosting Mobile AI Inference Performance by Removing Constraints on Graph OptimizationZhiying Xu, Hongding Peng, Wei WangINFOCOM 2023 · 被引用 1 次
- Mistify: Automating DNN Model Porting for On-Device Inference at the EdgePeizhen Guo, Bo Hu, Wenjun HuNSDI 2021 · 被引用 69 次
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM InferenceJiangwen Dong, Jiayu Li, Tianhang Zheng, Wanyu LINICML 2026 · 被引用 3 次
