Accelerating End-Cloud Collaborative Inference via Near Bubble-Free Pipeline Optimization
Luyao Gao, Jianchun Liu, Hongli Xu, Sun Xu, Qianpiao Ma, Liusheng Huang
摘要
End-cloud collaboration offers a promising strategy to enhance the Quality of Service (QoS) in DNN inference by offloading portions of the inference workload from end devices to cloud servers. Despite the potential, the complex model architectures and dynamic network conditions will introduce numerous bubbles (i.e., idle waiting time) in pipeline execution, resulting in inefficient resource utilization and degraded QoS. To address these challenges, we introduce a novel framework named COACH, designed for near bubble-free pipeline collaborative inference, thereby achieving low inference latency and high system throughput. Initially, COACH employs an offline component that utilizes an efficient recursive divide-and-conquer algorithm to optimize both model partitioning and transmission quantization, aiming to minimize the occurrence of pipeline bubbles. Subsequently, the online component in COACH employs an adaptive quantization adjustment and a context-aware caching strategy to further stabilize pipeline execution. Specifically, COACH analyzes the correlation between intermediate data and label semantic centers in the cache, along with its influence on the quantization adjustment, thereby effectively accommodating network fluctuations. Our experiments demonstrate the efficacy of COACH in reducing inference latency and enhancing system throughput. Notably, while maintaining comparable accuracy, COACH achieves up to 1.7× faster inference and 2.1× higher system throughput than baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis 等MobiCom 2020 · 被引用 312 次
- Distributed Inference Acceleration with Adaptive DNN Partitioning and OffloadingThaha Mohammed, Carlee Joe-Wong, Rohit Babbar, Mario Di FrancescoINFOCOM 2020 · 被引用 213 次
- CLIO: enabling automatic compilation of deep learning pipelines across IoT and cloudJin Huang, Colin Samplawski, Deepak Ganesan, Benjamin M. Marlin 等MobiCom 2020 · 被引用 72 次
- Zero Bubble (Almost) Pipeline ParallelismPenghui Qi, Xinyi Wan, Guangxing Huang, Min LinICLR 2024 · 被引用 31 次
- Boosting Mobile CNN Inference through Semantic MemoryYun Li, Chen Zhang, Shihao Han, Li Lyna Zhang 等ACM MM 2021 · 被引用 13 次
相关 Paper
- QoS-Aware Irregular Collaborative Inference for Improving Throughput of DNN ServicesKaihua Fu, Jiuchen Shi, Quan Chen, Ningxin Zheng 等SC 2022 · 被引用 7 次
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo 等UbiComp 2020 · 被引用 77 次
- Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge ServersZiyan Fu, Ju Ren, Deyu Zhang, Yuezhi Zhou 等INFOCOM 2022 · 被引用 28 次
- Concerto: Client-server Orchestration for Real-Time Video AnalyticsChaoyang Li, Rui-Xiao Zhang, Tianchi Huang, Lianchen Jia 等ACM MM 2023 · 被引用 8 次
- DynaPipe: Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline ParallelismHongxin Xu, Tianyu Guo, Xianwei ZhangNeurIPS 2025 · 被引用 4 次
