Distributed Inference Acceleration with Adaptive DNN Partitioning and Offloading
Thaha Mohammed, Carlee Joe-Wong, Rohit Babbar, Mario Di Francesco
Abstract
Deep neural networks (DNN) are the de-facto solution behind many intelligent applications of today, ranging from machine translation to autonomous driving. DNNs are accurate but resource-intensive, especially for embedded devices such as mobile phones and smart objects in the Internet of Things. To overcome the related resource constraints, DNN inference is generally offloaded to the edge or to the cloud. This is accomplished by partitioning the DNN and distributing computations at the two different ends. However, most of existing solutions simply split the DNN into two parts, one running locally or at the edge, and the other one in the cloud. In contrast, this article proposes a technique to divide a DNN in multiple partitions that can be processed locally by end devices or offloaded to one or multiple powerful nodes, such as in fog networks. The proposed scheme includes both an adaptive DNN partitioning scheme and a distributed algorithm to offload computations based on a matching game approach. Results obtained by using a self-driving car dataset and several DNN benchmarks show that the proposed solution significantly reduces the total latency for DNN inference compared to other distributed approaches and is 2.6 to 4.2 times faster than the state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 312781bd-ac74-4688-9940-d293ce6c09e0Cited by top-tier papers3
- Accelerating End-Cloud Collaborative Inference via Near Bubble-Free Pipeline OptimizationLuyao Gao, Jianchun Liu, Hongli Xu, Sun Xu et al.INFOCOM 2025 · 4 citations
- PhyDNNs: Bringing Deep Neural Networks to the Physical LayerMohammad Abdi, Khandaker Foysal Haque, Francesca Meneghello, Jonathan D. Ashdown et al.INFOCOM 2025 · 4 citations
- Efficient Distributed Inference of Deep Neural Networks via Restructuring and PruningAfshin Abdi, Saeed Rashidi, Faramarz Fekri, Tushar KrishnaAAAI 2023 · 3 citations
Related papers
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo et al.UbiComp 2020 · 77 citations
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia et al.MobiCom 2021 · 171 citations
- Resource-aware Deployment of Dynamic DNNs over Multi-tiered Interconnected SystemsChetna Singhal, Yashuo Wu, Francesco Malandrino, Marco Levorato et al.INFOCOM 2024 · 15 citations
- MTL-Split: Multi-Task Learning for Edge Devices using Split ComputingLuigi Capogrosso, Enrico Fraccaroli, Samarjit Chakraborty, Franco Fummi et al.DAC 2024 · 12 citations
- AoDNN: An Auto-Offloading Approach to Optimize Deep Inference for Fostering Mobile WebYakun Huang, Xiuquan Qiao, Schahram Dustdar, Yan LiINFOCOM 2022 · 17 citations
