Handling Network Faults in Distributed AI Training: Failover is Now an Option
Xin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui, Meng Wang, Yuze Jin, Pengfei Huo, Dongyang Wang, Lulu Chen, Lei Wang, Liaoyuan Feng, Xiaodong Liu
摘要
Distributed AI training often suffers from network faults. Network faults, especially at the last hop between a switch and a host, result in loss of connectivity, resulting in training job stalls and eventual failure. This is typically managed through a fail-stop mechanism, followed by a restart, incurring significant inefficiencies. We present ReCCL, the first network fault-tolerant collective communication library (CCL) that allows training progress to be preserved by seamlessly failing over to alternate paths when a network fault occurs. During failover, ReCCL keeps communication states synchronized while using dynamic channel load balancing and intra-host GPU routing to improve communication performance. Our evaluations demonstrate that ReCCL can perform failover seamlessly with minimal performance losses. Additionally, our simulations also demonstrate that failover can be effectively used to achieve significant savings in GPU hours for large-scale distributed AI training workloads.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- PReCCL: Performant and Resilient Collective Communication via Integrated Inband Telemetry and Workload ReallocationZhiyong Chen, Kaihui Gao, Li Chen, Rui Yan 等SIGCOMM 2026
- ResCCL: Resource-Efficient Scheduling for Collective CommunicationTongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao 等SIGCOMM 2025 · 被引用 11 次
- Comprehensive Deadlock Prevention for GPU Collective CommunicationLichen Pan, Juncheng Liu, Yongquan Fu, Jinhui Yuan 等EuroSys 2025 · 被引用 4 次
- TCCL: Discovering Better Communication Paths for PCIe GPU ClustersHeehoon Kim, Junyeol Ryu, Jaejin LeeASPLOS 2024 · 被引用 26 次
- CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingYida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun 等PPoPP 2026
