Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing
Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, Jaehyuk Huh
Abstract
As machine learning (ML) techniques are applied to a widening range of applications, high throughput ML inference serving has become critical for online services. Such ML inference servers with multiple GPUs pose new challenges in the scheduler design. First, they must provide a bounded latency for each request to support a consistent service-level objective (SLO). Second, they must be able to serve multiple heterogeneous ML models in a system, as cloud-based consolidation improves system utilization. To address the two requirements of ML inference servers, this paper proposes a new inference scheduling framework for multi-model ML inference servers. The paper shows that with SLO constraints, GPUs with growing parallelism are not fully utilized for ML inference tasks. To maximize the resource efficiency of GPUs, a key mechanism proposed in this paper is to exploit hardware support for spatial partitioning of GPU resources. With spatio-temporal sharing, a new abstraction layer of GPU resources is created with configurable GPU resources. The scheduler assigns requests to virtual GPUs, called gpulets, with the most effective amount of resources. The scheduler explores the three-dimensional search space with different batch sizes, temporal sharing, and spatial sharing efficiently. To minimize the cost for cloud-based inference servers, the framework auto-scales the required number of GPUs for a given workload. To consider the potential interference overheads when two ML tasks are running concurrently by spatially sharing a GPU, the scheduling decision is made with an interference prediction model. Our prototype implementation proves that the proposed spatio-temporal scheduling enhances throughput by 61.7% on average compared to the prior temporal scheduler, while satisfying SLOs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers36
- S3: Increasing GPU Utilization during Generative Inference for Higher ThroughputYunho Jin, Chun-Feng Wu, David Brooks, Gu-Yeon WeiNeurIPS 2023 · 150 citations
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete et al.OSDI 2024 · 125 citations
- Orion: Interference-aware, Fine-grained GPU Sharing for ML ApplicationsFoteini Strati, Xianzhe Ma, Ana KlimovicEuroSys 2024 · 96 citations
- Clover: Toward Sustainable AI with Carbon-Aware Machine Learning Inference ServiceBaolin Li, Siddharth Samsi, Vijay Gadepally, Devesh TiwariSC 2023 · 63 citations
- MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM ServingJiangfei Duan, Runyu Lu, Haojie Duanmu, Xiuhong Li et al.ICML 2024 · 51 citations
Builds on13
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 325 citations
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 184 citations
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 150 citations
- Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural NetworksSoroush Ghodrati, Byung Hoon Ahn, Joon Kyung Kim, Sean Kinzer et al.MICRO 2020 · 120 citations
Related papers
- BOER: Enhancing Resource Utilization for Deep Learning Inference with Hybrid Spatial GPU SharingBowen Zhang, Yuhang Wang, Zhuozhao LiSC 2025 · 4 citations
- KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference ServersMarcus Chow, Ali Jahanshahi, Daniel WongHPCA 2023 · 25 citations
- ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud EnvironmentsMunkyu Lee, Sihoon Seong, Minki Kang, Jihyuk Lee et al.SC 2024 · 19 citations
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU MultiplexingLei Gao, Chaoyi Jiang, Hossein Entezari Zarch, Daniel Wong et al.ICML 2026 · 4 citations
- Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye et al.EuroSys 2025 · 13 citations
