USENIX ATC2022顶会
PilotFish: Harvesting Free Cycles of Cloud Gaming with Deep Learning Training
Wei Zhang, Binghao Chen, Zhenhua Han, Quan Chen, Peng Cheng, Fan Yang, Ran Shu, Yuqing Yang, Minyi Guo
摘要
Cloud gaming services have become important workloads in cloud datacenter. However, our investigation shows that a cloud gaming service cannot saturate the modern cloud GPUs. One way to improve the GPU utilization is to co-locate multiple workloads within one GPU, which is challenging for cloud gaming due to its highly fluctuated and unpredictable GPU usage pattern. In this paper, we present PilotFish, a high-performance system that harvests the free GPU cycles of cloud gaming with deep learning (DL) training, while incurring almost zero interference to cloud gaming. We co-locate DL training jobs with cloud gaming, because they have stable and predictable workloads and have no strict latency requirement. In more detail, PilotFish captures the idle periods of the game's GPU usage with its low-overhead instrumentation to graphic libraries in sub-millisecond granularity. To avoid the potential interference to cloud gaming, PilotFish schedules training computation kernels only when they can finish before the idle GPU periods, and preempts straggler kernels running longer than expected. Our evaluation on popular cloud games and DL models shows PilotFish can harvest up to 85.1% of the idle GPU time from cloud gaming with no interference.
- This work is done while Wei Zhang is an intern in Microsoft Research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye 等EuroSys 2025 · 被引用 13 次
- SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUsYongkang Zhang, Haoxuan Yu, Chenxia Han, Cheng Wang 等PPoPP 2025 · 被引用 9 次
- High-density Mobile Cloud Gaming on Edge SoC ClustersLi Zhang, Shangguang Wang, Mengwei XuUSENIX ATC 2024 · 被引用 6 次
- SoCFlow: Efficient and Scalable DNN Training on SoC-Clustered Edge ServersDaliang Xu, Mengwei Xu, Chiheng Lou, Li Zhang 等ASPLOS 2024 · 被引用 5 次
它引用的顶会 Paper3
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang 等OSDI 2020 · 被引用 260 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- A Benchmarking Framework for Interactive 3D Applications in the CloudTianyi Liu, Sen He, Sunzhou Huang, Danny H. K. Tsang 等MICRO 2020 · 被引用 14 次
相关 Paper
- Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU ClustersQianhao Wu, Jiazhi Jiang, Guihui Ling, Yue PangVLDB 2026
- An Empirical Study on Low GPU Utilization of Deep Learning JobsYanjie Gao, Yichen He, Xinze Li, Bo Zhao 等ICSE 2024 · 被引用 22 次
- Zico: Efficient GPU Memory Sharing for Concurrent DNN TrainingGangmuk Lim, Jeongseob Ahn, Wencong Xiao, Youngjin Kwon 等USENIX ATC 2021 · 被引用 65 次
- Orion: Interference-aware, Fine-grained GPU Sharing for ML ApplicationsFoteini Strati, Xianzhe Ma, Ana KlimovicEuroSys 2024 · 被引用 96 次
- PipeSwitch: Fast Pipelined Context Switching for Deep Learning ApplicationsZhihao Bai, Zhen Zhang, Yibo Zhu, Xin JinOSDI 2020 · 被引用 152 次
