PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications
Zhihao Bai, Zhen Zhang, Yibo Zhu, Xin Jin
摘要
Deep learning (DL) workloads include throughputintensive training tasks and latency-sensitive inference tasks. The dominant practice today is to provision dedicated GPU clusters for training and inference separately. Due to the need to meet strict Service-Level Objectives (SLOs), GPU clusters are often over-provisioned based on the peak load with limited sharing between applications and task types.
We present PipeSwitch, a system that enables unused cycles of an inference application to be filled by training or other inference applications. It allows multiple DL applications to time-share the same GPU with the entire GPU memory and millisecond-scale switching overhead. With PipeSwitch, GPU utilization can be significantly improved without sacrificing SLOs. We achieve so by introducing pipelined context switching. The key idea is to leverage the layered structure of neural network models and their layer-by-layer computation pattern to pipeline model transmission over the PCIe and task execution in the GPU with model-aware grouping. We also design unified memory management and active-standby worker switching mechanisms to accompany the pipelining and ensure process-level isolation. We have built a PipeSwitch prototype and integrated it with PyTorch. Experiments on a variety of DL models and GPU cards show that PipeSwitch only incurs a task startup overhead of 3.6-6.6 ms and a total overhead of 5.4-34.6 ms (10-50× better than NVIDIA MPS), and achieves near 100% GPU utilization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper45
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu 等OSDI 2023 · 被引用 211 次
- SHEPHERD: Serving DNNs in the WildHong Zhang, Yupeng Tang, Anurag Khandelwal, Ion StoicaNSDI 2023 · 被引用 161 次
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 被引用 153 次
- Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNsJohn Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao 等NSDI 2023 · 被引用 144 次
- Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient DescentQizhen Weng, Lingyun Yang, Yinghao Yu, Wei Wang 等USENIX ATC 2023 · 被引用 115 次
它引用的顶会 Paper3
- SwapAdvisor: Pushing Deep Learning Beyond the GPU Memory Limit via Smart SwappingChien-Chin Huang, Gu Jin, Jinyang LiASPLOS 2020 · 被引用 161 次
- Themis: Fair and Efficient GPU Cluster SchedulingKshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman 等NSDI 2020 · 被引用 22 次
- On Efficient Constructions of CheckpointsYu Chen, Zhenming Liu, Bin Ren, Xin JinICML 2020 · 被引用 21 次
相关 Paper
- Transparent GPU Sharing in Container Clouds for Deep Learning WorkloadsBingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu 等NSDI 2023 · 被引用 112 次
- PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline ParallelismZ. Jonny Kong, Qiang Xu, Y. Charlie HuUSENIX ATC 2025 · 被引用 4 次
- Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye 等EuroSys 2025 · 被引用 13 次
- DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUsAmir Fakhim Babaei, Thidapat ChantemDAC 2025 · 被引用 4 次
- Orion: Interference-aware, Fine-grained GPU Sharing for ML ApplicationsFoteini Strati, Xianzhe Ma, Ana KlimovicEuroSys 2024 · 被引用 96 次
