Accelerating End-Cloud Collaborative Inference via Near Bubble-Free Pipeline Optimization
Luyao Gao, Jianchun Liu, Hongli Xu, Sun Xu, Qianpiao Ma, Liusheng Huang
Abstract
End-cloud collaboration offers a promising strategy to enhance the Quality of Service (QoS) in DNN inference by offloading portions of the inference workload from end devices to cloud servers. Despite the potential, the complex model architectures and dynamic network conditions will introduce numerous bubbles (i.e., idle waiting time) in pipeline execution, resulting in inefficient resource utilization and degraded QoS. To address these challenges, we introduce a novel framework named COACH, designed for near bubble-free pipeline collaborative inference, thereby achieving low inference latency and high system throughput. Initially, COACH employs an offline component that utilizes an efficient recursive divide-and-conquer algorithm to optimize both model partitioning and transmission quantization, aiming to minimize the occurrence of pipeline bubbles. Subsequently, the online component in COACH employs an adaptive quantization adjustment and a context-aware caching strategy to further stabilize pipeline execution. Specifically, COACH analyzes the correlation between intermediate data and label semantic centers in the cache, along with its influence on the quantization adjustment, thereby effectively accommodating network fluctuations. Our experiments demonstrate the efficacy of COACH in reducing inference latency and enhancing system throughput. Notably, while maintaining comparable accuracy, COACH achieves up to 1.7× faster inference and 2.1× higher system throughput than baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis et al.MobiCom 2020 · 312 citations
- Distributed Inference Acceleration with Adaptive DNN Partitioning and OffloadingThaha Mohammed, Carlee Joe-Wong, Rohit Babbar, Mario Di FrancescoINFOCOM 2020 · 213 citations
- CLIO: enabling automatic compilation of deep learning pipelines across IoT and cloudJin Huang, Colin Samplawski, Deepak Ganesan, Benjamin M. Marlin et al.MobiCom 2020 · 72 citations
- Zero Bubble (Almost) Pipeline ParallelismPenghui Qi, Xinyi Wan, Guangxing Huang, Min LinICLR 2024 · 31 citations
- Boosting Mobile CNN Inference through Semantic MemoryYun Li, Chen Zhang, Shihao Han, Li Lyna Zhang et al.ACM MM 2021 · 13 citations
Related papers
- QoS-Aware Irregular Collaborative Inference for Improving Throughput of DNN ServicesKaihua Fu, Jiuchen Shi, Quan Chen, Ningxin Zheng et al.SC 2022 · 7 citations
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo et al.UbiComp 2020 · 77 citations
- Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge ServersZiyan Fu, Ju Ren, Deyu Zhang, Yuezhi Zhou et al.INFOCOM 2022 · 28 citations
- Concerto: Client-server Orchestration for Real-Time Video AnalyticsChaoyang Li, Rui-Xiao Zhang, Tianchi Huang, Lianchen Jia et al.ACM MM 2023 · 8 citations
- DynaPipe: Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline ParallelismHongxin Xu, Tianyu Guo, Xianwei ZhangNeurIPS 2025 · 4 citations
