Distributed Inference with Deep Learning Models across Heterogeneous Edge Devices
Chenghao Hu, Baochun Li
Abstract
Recent years witnessed an increasing research attention in deploying deep learning models on edge devices for inference. Due to limited capabilities and power constraints, it may be necessary to distribute the inference workload across multiple devices. Existing mechanisms divided the model across edge devices with the assumption that deep learning models are constructed with a chain of layers. In reality, however, modern deep learning models are more complex, involving a directed acyclic graph (DAG) rather than a chain of layers.
In this paper, we present EdgeFlow, a new distributed inference mechanism designed for general DAG structured deep learning models. Specifically, EdgeFlow partitions model layers into independent execution units with a new progressive model partitioning algorithm. By producing near-optimal model partitions, our new algorithm seeks to improve the run-time performance of distributed inference as these partitions are distributed across the edge devices. During inference, EdgeFlow orchestrates the intermediate results flowing through these units to fulfill the complicated layer dependencies. We have implemented Edge-Flow based on PyTorch, and evaluated it with state-of-theart deep learning models in different structures. The results show that EdgeFlow reducing the inference latency by up to 40.2% compared with other approaches, which demonstrates the effectiveness of our design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed7ab284-7f46-41e5-8785-1ef548697eacCited by top-tier papers3
- Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AIZhi Zhou, Jiajie Xie, Mengke Huang, Tao Ouyang et al.INFOCOM 2025 · 3 citations
- ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer InferenceXiao Liu, Lijun Zhang, Deepak Ganesan, Hui GuanICML 2026
- Hierarchical Inference with an Offload QueueSrinivas Nomula, Parimal Parag, Ayalvadi Ganesh, Jaya Prakash ChampatiINFOCOM 2026
Related papers
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo et al.UbiComp 2020 · 77 citations
- Resource-aware Deployment of Dynamic DNNs over Multi-tiered Interconnected SystemsChetna Singhal, Yashuo Wu, Francesco Malandrino, Marco Levorato et al.INFOCOM 2024 · 15 citations
- AGO: Boosting Mobile AI Inference Performance by Removing Constraints on Graph OptimizationZhiying Xu, Hongding Peng, Wei WangINFOCOM 2023 · 1 citation
- Mistify: Automating DNN Model Porting for On-Device Inference at the EdgePeizhen Guo, Bo Hu, Wenjun HuNSDI 2021 · 69 citations
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM InferenceJiangwen Dong, Jiayu Li, Tianhang Zheng, Wanyu LINICML 2026 · 3 citations
