GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at Scale
Tianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang, Wenchao Wu, Qinkai Duan, Guodong Yang, Jiamang Wang, Lin Qu, Liping Zhang
Abstract
Fail-slows, or stragglers, are common problems in large-scale hybrid-parallel training that runs on a large fleet of GPU servers for an extended period of time. Yet, these problems are not well studied. In this paper, we first present a characterization study on a shared production cluster with over 10,000 GPUs. We find that fail-slows manifest as transient stragglers caused by slow computations or communications due to contention, device degradation, or network congestion, lasting from sub-minutes to nearly ten hours, and delaying large training jobs by 1.34× on average. The current practice is to manually detect fail-slows and treat them as fail-stops by means of checkpoint-and-restart failover, which is time-consuming. In this paper, we propose GREYHOUND, a system that rapidly identifies slow GPUs and/or communication links, and effectively tackles them with a novel multi-level mitigation mechanism, all without human intervention. GREYHOUND correctly detects fail-slows in a production cluster with over 99% accuracy. Testbed experiment on 256 H800 GPUs further shows it effectively handles (manually injected) stragglers, improving end-to-end throughput by 1.58×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eea695c9-96f3-4d44-b05e-868dcd7e0673Cited by top-tier papers8
- FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus ScaleWeihao Cui, Ji Zhang, Han Zhao, Chao Liu et al.NSDI 2026 · 7 citations
- Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingYangtao Deng, Lei Zhang, Qinlong Wang, Xiaoyun Zhi et al.SOSP 2025 · 7 citations
- CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingYida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun et al.PPoPP 2026
- RobustRL: Role-Based Fault Tolerance System for RL Post-TrainingZhenqian Chen, Baoquan Zhong, Xiang Li, Qing Dai et al.OSDI 2026
- ResiHP: Taming LLM Training Failures with Dynamic Hybrid ParallelismTenghui Ma, Jihu Guo, Wei Gao, Sitian Lu et al.HPDC 2026
Builds on23
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase et al.USENIX ATC 2021 · 657 citations
Related papers
- Holmes: Localizing Irregularities in LLM Training with Mega-scale GPU ClustersZhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia et al.NSDI 2025 · 23 citations
- Understanding Stragglers in Large Model Training Using What-if AnalysisJinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao et al.OSDI 2025 · 23 citations
- Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model ParallelizationHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin et al.SIGMOD 2025 · 6 citations
- EROICA: Online Performance Troubleshooting for Large-scale Model TrainingYu Guan, Zhiyu Yin, Haoyu Chen, Sheng Cheng et al.NSDI 2026 · 1 citation
- Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed ClustersFoteini Strati, Zhendong Zhang, George Manos, Ixeia Sánchez Périz et al.SOSP 2025 · 2 citations
