BOER: Enhancing Resource Utilization for Deep Learning Inference with Hybrid Spatial GPU Sharing
Bowen Zhang, Yuhang Wang, Zhuozhao Li
摘要
Many inference systems leverage spatial multiplexing technologies, such as Multi-Process Service (MPS) and Multi-Instance GPU (MIG), to serve deep learning models concurrently on a single GPU. However, existing solutions suffer from interference under MPS and rigid partition sizes in MIG. To address these limitations, we propose BOER, a system that combines MPS atop MIG partitions to reduce interference and enhance GPU utilization. BOER identifies key challenges in integrating MPS with MIG and introduces a hierarchical scheduling framework that jointly determines model colocation, workload distribution, MIG partitioning, and MPS configurations, while minimizing resource fragmentation and MIG reconfiguration overhead. Since MPS interference is difficult to predict accurately, BOER avoids performance models and instead employs a Bayesian optimization with tailored acceleration strategies to efficiently explore the MPS configuration space. Evaluation on a real testbed demonstrates that BOER outperforms state-of-the-art spatial multiplexing solutions, improving inference throughput by up to 46.04%–77.19% while preserving Quality-of-Service requirements.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud EnvironmentsMunkyu Lee, Sihoon Seong, Minki Kang, Jihyuk Lee 等SC 2024 · 被引用 19 次
- KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference ServersMarcus Chow, Ali Jahanshahi, Daniel WongHPCA 2023 · 被引用 25 次
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 被引用 35 次
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU MultiplexingLei Gao, Chaoyi Jiang, Hossein Entezari Zarch, Daniel Wong 等ICML 2026 · 被引用 4 次
