Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU Clusters
Qianhao Wu, Jiazhi Jiang, Guihui Ling, Yue Pang
Abstract
Online deep learning (DL) training has become pivotal in powering real-time applications. Yet tidal workload fluctuations leave GPU clusters significantly underutilized during off-peak periods. This imbalance not only wastes GPU capacity but also exacerbates scarcity for other GPU-intensive jobs on cloud-native GPU cluster. Cluster-wide resource leasing across different tenants enabled by elastic scaling offers a promising opportunity to enhance GPU utilization for cloud-native online DL training deployments in multi-tenant GPU clusters. Existing solutions do not address the unique challenges associated with maintaining system stability during elastic scaling for online DL training jobs, including prolonged disruptions due to job reconstruction, failures arising from dependency-unaware operation triggering, and the unreliable reclamation of high-availability GPU resources. In this paper, we introduce WeFlex, a resilience-aware elastic scaling solution engineered for cloud-native online deep learning jobs in multi-tenant GPU clusters. WeFlex enables online training jobs to lease idle GPUs for other GPU-intensive jobs during low-demand periods while ensuring rapid reclamation as demand surges. It significantly reduces the duration of training disruptions through constructing a interruption mitigation pipeline, prevents dependency-unaware operation failures via topology-aware pod orchestration, and ensure reclamation of high-availability GPU resources through right-of-return GPU leasing. Evaluations on 10,000-plus scale GPU clusters in production demonstrate that WeFlex significantly enhances GPU utilization while reliably maintaining continuous training performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0c0dfe7-cc3b-44eb-92db-0302689645a0Builds on8
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger et al.OSDI 2021 · 258 citations
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- Looking Beyond GPUs for DNN Scheduling on Multi-Tenant ClustersJayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay ChidambaramOSDI 2022 · 91 citations
- ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep LearningDiandian Gu, Yihao Zhao, Yinmin Zhong, Yifan Xiong et al.ASPLOS 2023 · 62 citations
Related papers
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun et al.SC 2023 · 16 citations
- An efficient and non-intrusive GPU scheduling framework for deep learning training systemsShaoqi Wang, Oscar J. Gonzalez, Xiaobo Zhou, Thomas Williams et al.SC 2020 · 21 citations
- Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective ElasticityCunchi Lv, Xiao Shi, Zhengyu Lei, Jinyue Huang et al.ASPLOS 2025 · 10 citations
- Lyra: Elastic Scheduling for Deep Learning ClustersJiamin Li, Hong Xu, Yibo Zhu, Zherui Liu et al.EuroSys 2023 · 59 citations
- Online evolutionary batch size orchestration for scheduling deep learning workloads in GPU clustersZhengda Bian, Shenggui Li, Wei Wang, Yang YouSC 2021 · 22 citations
