USENIX ATC2025顶会
Barre: Empowering Simplified and Versatile Programmable Congestion Control in High-Speed AI Clusters
Yajuan Peng, Haoran Wei, Xiaolong Zhong, Junkai Huang, Haohan Xu, Zicheng Wang, Yang Bai, Zhuo Jiang, Jianxi Ye, Xiaoliang Wang, Xiaoming Fu, Huichen Dai
摘要
Network interface cards (NICs) and switches have entered the 400 Gbps era. RoCEv2 networks face significant challenges in congestion management, particularly under high-throughput workloads. While advanced congestion control algorithms have been proposed, their deployment in large-scale data centers remains hindered by complex parameter tuning and dependency on sophisticated hardware features. In this paper, we present Barre, a simple yet highly effective congestion control scheme designed for modern AI/HPC clusters operating at 400 Gbps. By leveraging commodity hardware and standard network functionalities, Barre achieves near-optimal performance in fairness, congestion responsiveness, and scalability with minimal overhead. Deployed in our 400 Gbps RoCE cluster for over a year and supporting up to 10,000 GPUs, Barre improves AI training task throughput by an average of 9.6%. Furthermore, we demonstrate that Barre's core principles can be seamlessly applied to enhance DCQCN, a widely deployed congestion control algorithm, underscoring its practicality and versatility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 等NSDI 2024 · 被引用 415 次
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel 等SIGCOMM 2020 · 被引用 333 次
- Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao 等SIGCOMM 2024 · 被引用 173 次
- RDMA over Ethernet for Distributed Training at Meta ScaleAdithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu 等SIGCOMM 2024 · 被引用 171 次
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng 等NSDI 2023 · 被引用 154 次
相关 Paper
- PC4: Precision Collective Communication Congestion Control for AI ClusterTaoran Qi, Shuo Li, Xingqi Zou, Liangce Deng 等INFOCOM 2025 · 被引用 3 次
- Backpressure Flow ControlPrateesh Goyal, Preey Shah, Kevin Zhao, Georgios Nikolaidis 等NSDI 2022
- High-throughput and Flexible Host Networking for Accelerated ComputingAthinagoras Skiadopoulos, Zhiqiang Xie, Mark Zhao, Qizhe Cai 等OSDI 2024 · 被引用 11 次
- STORM: Enabling Traffic Scheduling for RDMAJichun Wu, Ran Shu, Gianni Antichi, Yongqiang Xiong 等SIGCOMM 2026
- BCC: Re-architecting Congestion Control in DCNsQingkai Meng, Shan Zhang, Zhiyuan Wang, Tao Tong 等INFOCOM 2024 · 被引用 11 次
