SGLB: Scalable and Robust Global Load Balancing in Commodity AI Clusters
Chenchen Qi, Wenfei Wu, Yongcan Wang, Keqiang He, Yu-Hsiang Kao, Zongying He, Chen-Yu Yen, Zhuo Jiang, Feng Luo, Surendra Anubolu, Yanjin Gao, Bingfeng Lin
摘要
Internet companies are constructing large-scale AI clusters with commodity Ethernet switches for AI model training to support their businesses. AI training workloads impose stringent network requirements, mandating that cluster networks deliver high peak throughput while maintaining robustness and resilience in the face of link failures. We present SGLB, a distributed, global congestion-aware load balancing system for AI clusters. SGLB operates a control-plane protocol, SyncMesh, to enable a new load balancing abstraction in modern commodity switches—Global Load Balancing (GLB) engine—which utilizes global congestion information to distribute traffic across all available paths. We address three key challenges in designing SGLB: fast routing convergence to minimize downtime in the event of link failures, scalable maintenance of congestion profiles within the constraints of limited switch hardware resources, and preventing GLB throughput suppression in scenarios where path bandwidths are asymmetric. We prototype SGLB and conduct extensive experiments to evaluate SGLB. SGLB ensures rapid routing convergence in the event of link failures, recovering in as little as 45 μs to guarantee network robustness for long-term, stable model training. Additionally, SGLB effectively load-balances traffic across paths, avoiding those with global congestion, which accelerates All-to-All collective communication by up to 60%.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- EPIC: Abstraction and Polymorphism of In-Network Collectives on EthernetYitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou 等SIGCOMM 2026 · 被引用 1 次
- MonkeyTree: Near-Minimal Congestion for Multi-tenant Training via MigrationAnton A. Zabreyko, Weiyang Wang, Manya GhobadiSIGCOMM 2026
相关 Paper
- PC4: Precision Collective Communication Congestion Control for AI ClusterTaoran Qi, Shuo Li, Xingqi Zou, Liangce Deng 等INFOCOM 2025 · 被引用 3 次
- Handling Network Faults in Distributed AI Training: Failover is Now an OptionXin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui 等EuroSys 2026 · 被引用 1 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- AdaGL: Adaptive Learning for Agile Distributed Training of Gigantic GNNsRuisi Zhang, Mojan Javaheripi, Zahra Ghodsi, Amit Bleiweiss 等DAC 2023 · 被引用 3 次
- FAST: An Efficient Scheduler for All-to-All GPU CommunicationYiran Lei, Dongjoo Lee, Liangyu Zhao, Daniar Kurniawan 等NSDI 2026 · 被引用 13 次
