SC2025Top-tier venue
BOER: Enhancing Resource Utilization for Deep Learning Inference with Hybrid Spatial GPU Sharing
Bowen Zhang, Yuhang Wang, Zhuozhao Li
Abstract
Many inference systems leverage spatial multiplexing technologies, such as Multi-Process Service (MPS) and Multi-Instance GPU (MIG), to serve deep learning models concurrently on a single GPU. However, existing solutions suffer from interference under MPS and rigid partition sizes in MIG. To address these limitations, we propose BOER, a system that combines MPS atop MIG partitions to reduce interference and enhance GPU utilization. BOER identifies key challenges in integrating MPS with MIG and introduces a hierarchical scheduling framework that jointly determines model colocation, workload distribution, MIG partitioning, and MPS configurations, while minimizing resource fragmentation and MIG reconfiguration overhead. Since MPS interference is difficult to predict accurately, BOER avoids performance models and instead employs a Bayesian optimization with tailored acceleration strategies to efficiently explore the MPS configuration space. Evaluation on a real testbed demonstrates that BOER outperforms state-of-the-art spatial multiplexing solutions, improving inference throughput by up to 46.04%–77.19% while preserving Quality-of-Service requirements.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 25741379-5289-4243-90cb-5833407fb9c1Cited by top-tier papers1
Ask how each one uses itRelated papers
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud EnvironmentsMunkyu Lee, Sihoon Seong, Minki Kang, Jihyuk Lee et al.SC 2024 · 19 citations
- KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference ServersMarcus Chow, Ali Jahanshahi, Daniel WongHPCA 2023 · 25 citations
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 35 citations
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU MultiplexingLei Gao, Chaoyi Jiang, Hossein Entezari Zarch, Daniel Wong et al.ICML 2026 · 4 citations
