SGLB: Scalable and Robust Global Load Balancing in Commodity AI Clusters
Chenchen Qi, Wenfei Wu, Yongcan Wang, Keqiang He, Yu-Hsiang Kao, Zongying He, Chen-Yu Yen, Zhuo Jiang, Feng Luo, Surendra Anubolu, Yanjin Gao, Bingfeng Lin
Abstract
Internet companies are constructing large-scale AI clusters with commodity Ethernet switches for AI model training to support their businesses. AI training workloads impose stringent network requirements, mandating that cluster networks deliver high peak throughput while maintaining robustness and resilience in the face of link failures. We present SGLB, a distributed, global congestion-aware load balancing system for AI clusters. SGLB operates a control-plane protocol, SyncMesh, to enable a new load balancing abstraction in modern commodity switches—Global Load Balancing (GLB) engine—which utilizes global congestion information to distribute traffic across all available paths. We address three key challenges in designing SGLB: fast routing convergence to minimize downtime in the event of link failures, scalable maintenance of congestion profiles within the constraints of limited switch hardware resources, and preventing GLB throughput suppression in scenarios where path bandwidths are asymmetric. We prototype SGLB and conduct extensive experiments to evaluate SGLB. SGLB ensures rapid routing convergence in the event of link failures, recovering in as little as 45 μs to guarantee network robustness for long-term, stable model training. Additionally, SGLB effectively load-balances traffic across paths, avoiding those with global congestion, which accelerates All-to-All collective communication by up to 60%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 83354593-46f1-4ac8-9c86-116d61a4dd95Cited by top-tier papers2
- EPIC: Abstraction and Polymorphism of In-Network Collectives on EthernetYitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou et al.SIGCOMM 2026 · 1 citation
- MonkeyTree: Near-Minimal Congestion for Multi-tenant Training via MigrationAnton A. Zabreyko, Weiyang Wang, Manya GhobadiSIGCOMM 2026
Related papers
- PC4: Precision Collective Communication Congestion Control for AI ClusterTaoran Qi, Shuo Li, Xingqi Zou, Liangce Deng et al.INFOCOM 2025 · 3 citations
- Handling Network Faults in Distributed AI Training: Failover is Now an OptionXin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui et al.EuroSys 2026 · 1 citation
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- AdaGL: Adaptive Learning for Agile Distributed Training of Gigantic GNNsRuisi Zhang, Mojan Javaheripi, Zahra Ghodsi, Amit Bleiweiss et al.DAC 2023 · 3 citations
- FAST: An Efficient Scheduler for All-to-All GPU CommunicationYiran Lei, Dongjoo Lee, Liangyu Zhao, Daniar Kurniawan et al.NSDI 2026 · 13 citations
