EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUs
Mingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun, Hanyu Zhao, Shiru Ren, Zhongzhi Luan, Xianyan Jia, Yi Liu, Yong Li, Wei Lin, Depei Qian
摘要
Distributed synchronized GPU training is commonly used for deep learning. The resource constraint of using a fixed number of GPUs makes large-scale training jobs suffer from long queuing time for resource allocation, and lowers the cluster utilization. Adapting to resource elasticity can alleviate this but often introduces inconsistent model accuracy, due to lacking of capability to decouple model training procedure from resource allocation. We propose EasyScale, an elastic training system that achieves consistent model accuracy under resource elasticity for both homogeneous and heterogeneous GPUs. EasyScale preserves the data-parallel training behaviors strictly, traces the consistency-relevant factors carefully, utilizes the deep learning characteristics for EasyScaleThread abstraction and fast context-switching. To utilize heterogeneous cluster, EasyScale dynamically assigns workers based on the intra-/inter-job schedulers, minimizing load imbalance and maximizing aggregated job throughput. Deployed in an online serving cluster, EasyScale powers the training jobs to utilize idle GPUs opportunistically, improving overall cluster utilization by 62.1%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Metis: Fast Automatic Distributed Training on Heterogeneous GPUsTaegeon Um, Byungsoo Oh, Minyoung Kang, Woo-Yeon Lee 等USENIX ATC 2024 · 被引用 81 次
- ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model DevelopmentBorui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng 等NSDI 2025 · 被引用 46 次
- Colocating ML Inference and Training with Fast GPU Memory HandoverJiali Wang, Yankui Wang, Mingcong Han, Rong ChenUSENIX ATC 2025 · 被引用 12 次
- Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor CollectionsMarcel Wagenländer, Guo Li, Bo Zhao, Luo Mai 等SOSP 2024 · 被引用 8 次
- FlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length InputsHairui Zhao, Qi Tian, Hongliang Li, Zizhong ChenUSENIX ATC 2025 · 被引用 6 次
它引用的顶会 Paper20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee 等OSDI 2020 · 被引用 286 次
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang 等OSDI 2020 · 被引用 260 次
相关 Paper
- JABAS: Joint Adaptive Batching and Automatic Scaling for DNN Training on Heterogeneous GPUsGyeongchan Yun, Junesoo Kang, Hyunjoon Jeong, Sanghyeon Eom 等EuroSys 2025 · 被引用 2 次
- Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU ClustersQianhao Wu, Jiazhi Jiang, Guihui Ling, Yue PangVLDB 2026
- Hare: Exploiting Inter-job and Intra-job Parallelism of Distributed Machine Learning on Heterogeneous GPUsFahao Chen, Peng Li, Celimuge Wu, Song GuoHPDC 2022 · 被引用 10 次
- Lyra: Elastic Scheduling for Deep Learning ClustersJiamin Li, Hong Xu, Yibo Zhu, Zherui Liu 等EuroSys 2023 · 被引用 59 次
- ElasGNN: An Elastic Training Framework for Distributed GNN TrainingSiqi Wang, Hailong Yang, Pengbo Wang, Hongliang Cao 等PPoPP 2026
