Heet: Accelerating Elastic Training in Heterogeneous Deep Learning Clusters
Zizhao Mo, Huanle Xu, Chengzhong Xu
2024Year
22Citations
7Top-tier citations
Abstract
Modern GPU clusters inherently exhibit heterogeneity, encompassing various aspects such as computation and communication. This heterogeneity poses a significant challenge for the elastic scheduling of deep learning workloads. Unfortunately, existing elastic schedulers often overlook the impact of heterogeneity on scaling efficiency, resulting in considerably prolonged job completion times.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c02a89c8-89c7-461f-95e5-960338f0714aCited by top-tier papers7
- Metis: Fast Automatic Distributed Training on Heterogeneous GPUsTaegeon Um, Byungsoo Oh, Minyoung Kang, Woo-Yeon Lee et al.USENIX ATC 2024 · 81 citations
- Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model ParallelizationHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin et al.SIGMOD 2025 · 6 citations
- Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront SchedulingYujie Wang, Shenhan Zhu, Fangcheng Fu, Xupeng Miao et al.ASPLOS 2025 · 6 citations
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic ParallelismZizhao Mo, Jianxiong Liao, Huanle Xu, Zhi Zhou et al.SC 2025 · 3 citations
- Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed ClustersFoteini Strati, Zhendong Zhang, George Manos, Ixeia Sánchez Périz et al.SOSP 2025 · 2 citations
Related papers
- Sia: Heterogeneity-aware, goodput-optimized ML-cluster schedulingSuhas Jayaram Subramanya, Daiyaan Arfeen, Shouxu Lin, Aurick Qiao et al.SOSP 2023 · 50 citations
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun et al.SC 2023 · 16 citations
- Hare: Exploiting Inter-job and Intra-job Parallelism of Distributed Machine Learning on Heterogeneous GPUsFahao Chen, Peng Li, Celimuge Wu, Song GuoHPDC 2022 · 10 citations
- Online evolutionary batch size orchestration for scheduling deep learning workloads in GPU clustersZhengda Bian, Shenggui Li, Wei Wang, Yang YouSC 2021 · 22 citations
- An efficient and non-intrusive GPU scheduling framework for deep learning training systemsShaoqi Wang, Oscar J. Gonzalez, Xiaobo Zhou, Thomas Williams et al.SC 2020 · 21 citations
