Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU Clusters
Qianhao Wu, Jiazhi Jiang, Guihui Ling, Yue Pang
摘要
Online deep learning (DL) training has become pivotal in powering real-time applications. Yet tidal workload fluctuations leave GPU clusters significantly underutilized during off-peak periods. This imbalance not only wastes GPU capacity but also exacerbates scarcity for other GPU-intensive jobs on cloud-native GPU cluster. Cluster-wide resource leasing across different tenants enabled by elastic scaling offers a promising opportunity to enhance GPU utilization for cloud-native online DL training deployments in multi-tenant GPU clusters. Existing solutions do not address the unique challenges associated with maintaining system stability during elastic scaling for online DL training jobs, including prolonged disruptions due to job reconstruction, failures arising from dependency-unaware operation triggering, and the unreliable reclamation of high-availability GPU resources. In this paper, we introduce WeFlex, a resilience-aware elastic scaling solution engineered for cloud-native online deep learning jobs in multi-tenant GPU clusters. WeFlex enables online training jobs to lease idle GPUs for other GPU-intensive jobs during low-demand periods while ensuring rapid reclamation as demand surges. It significantly reduces the duration of training disruptions through constructing a interruption mitigation pipeline, prevents dependency-unaware operation failures via topology-aware pod orchestration, and ensure reclamation of high-availability GPU resources through right-of-return GPU leasing. Evaluations on 10,000-plus scale GPU clusters in production demonstrate that WeFlex significantly enhances GPU utilization while reliably maintaining continuous training performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang 等OSDI 2020 · 被引用 260 次
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger 等OSDI 2021 · 被引用 258 次
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- Looking Beyond GPUs for DNN Scheduling on Multi-Tenant ClustersJayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay ChidambaramOSDI 2022 · 被引用 91 次
- ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep LearningDiandian Gu, Yihao Zhao, Yinmin Zhong, Yifan Xiong 等ASPLOS 2023 · 被引用 62 次
相关 Paper
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun 等SC 2023 · 被引用 16 次
- An efficient and non-intrusive GPU scheduling framework for deep learning training systemsShaoqi Wang, Oscar J. Gonzalez, Xiaobo Zhou, Thomas Williams 等SC 2020 · 被引用 21 次
- Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective ElasticityCunchi Lv, Xiao Shi, Zhengyu Lei, Jinyue Huang 等ASPLOS 2025 · 被引用 10 次
- Lyra: Elastic Scheduling for Deep Learning ClustersJiamin Li, Hong Xu, Yibo Zhu, Zherui Liu 等EuroSys 2023 · 被引用 59 次
- Online evolutionary batch size orchestration for scheduling deep learning workloads in GPU clustersZhengda Bian, Shenggui Li, Wei Wang, Yang YouSC 2021 · 被引用 22 次
