KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container Cloud
Ting-An Yeh, Hung-Hsin Chen, Jerry Chou
摘要
Container has emerged as a new technology in clouds to replace virtual machines (VM) for distributed applications deployment and operation. With the increasing number of new cloud-focused applications, such as deep learning and high performance applications, started to reply on the high computing throughput of GPUs, efficiently supporting GPU in container cloud becomes essential. While GPU virtualization has been extensively studied for VM, limited work has been done for containers. One of the key challenges is the lack of support for GPU sharing between multiple concurrent containers. This limitation leads to low resource utilization when a GPU device cannot be fully utilized by a single application due to the burstiness of GPU workload and the limited memory bandwidth. To overcome this issue, we designed and implemented KubeShare, which extends Kubernetes to enable GPU sharing with fine-grained allocation. KubeShare is the first solution for Kubernetes to make GPU device as a first class resources for scheduling and allocations. Using real deep learning workloads, we demonstrated KubeShare can significantly increase GPU utilization and overall system throughput around 2x with less than 10% performance overhead during container initialization and execution.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient DescentQizhen Weng, Lingyun Yang, Yinghao Yu, Wei Wang 等USENIX ATC 2023 · 被引用 115 次
- Efficient Performance-Aware GPU Sharing with Compatibility and Isolation through Kernel Space InterceptionShulai Zhang, Ao Xu, Quan Chen, Han Zhao 等USENIX ATC 2025 · 被引用 16 次
- Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective ElasticityCunchi Lv, Xiao Shi, Zhengyu Lei, Jinyue Huang 等ASPLOS 2025 · 被引用 10 次
- XSched: Preemptive Scheduling for Diverse XPUsWeihang Shen, Mingcong Han, Jialong Liu, Rong Chen 等OSDI 2025 · 被引用 9 次
- Are We Ready for Vision-Centric Driving Streaming Perception? The ASAP BenchmarkXiaofeng Wang, Zheng Zhu, Yunpeng Zhang, Guan Huang 等CVPR 2023
相关 Paper
- Transparent GPU Sharing in Container Clouds for Deep Learning WorkloadsBingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu 等NSDI 2023 · 被引用 112 次
- An efficient and non-intrusive GPU scheduling framework for deep learning training systemsShaoqi Wang, Oscar J. Gonzalez, Xiaobo Zhou, Thomas Williams 等SC 2020 · 被引用 21 次
- Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning WorkloadsWei Zhao, Anand Jayarajan, Gennady PekhimenkoASPLOS 2025 · 被引用 4 次
- gShare: Efficient GPU Sharing with Aggressive Scheduling in Multi-tenant FaaS platformYanan Yang, Zhengxiong Jiang, Meiqi Zhu, Hongqiang Xu 等ASPLOS 2026 · 被引用 1 次
- Balancing efficiency and fairness in heterogeneous GPU clusters for deep learningShubham Chaudhary, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra 等EuroSys 2020 · 被引用 135 次
