Alibaba HPN: A Data Center Network for Large Language Model Training
Kun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao, Yichi Xu, Yu Guan, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao, Chao Wang, Peng Wang
摘要
This paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. This requires us to design a new data center network architecture specifically for LLM training. Unlike general cloud computing which generates millions of small flows (e.g., lower than 10Gbps), LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP), the commonly used load-balancing scheme in traditional data centers, to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization by decreasing the occurrences of ECMP, but also greatly reduces the search space for path selection, thus allowing us to precisely select network paths capable of holding elephant flows. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to single-point failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks by addressing the layer-2 synchronization challenges. HPN has been deployed in our production for more than eight months. We share our experience in motivating, designing, and building HPN, as well as the operational lessons of HPN in production.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu 等NSDI 2025 · 被引用 82 次
- White-Boxing RDMA with Packet-Granular Software ControlChenxingyu Zhao, Jaehong Min, Ming Liu, Arvind KrishnamurthyNSDI 2025 · 被引用 28 次
- Evolution of Aegis: Fault Diagnosis for AI Model Training Service in ProductionJianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng 等NSDI 2025 · 被引用 21 次
- Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep LearningWei An, Xiao Bi, Guanting Chen, Shanhuang Chen 等SC 2024 · 被引用 20 次
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang 等USENIX ATC 2025 · 被引用 19 次
它引用的顶会 Paper6
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh 等SIGCOMM 2022 · 被引用 230 次
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre 等NSDI 2023 · 被引用 117 次
- PLB: congestion signals are simple and effective for network load balancingMubashir Adnan Qureshi, Yuchung Cheng, Qianwen Yin, Qiaobin Fu 等SIGCOMM 2022 · 被引用 82 次
- Lightwave Fabrics: At-Scale Optical Circuit Switching for Datacenter and Machine Learning SystemsHong Liu, Ryohei Urata, Kevin Yasumura, Xiang Zhou 等SIGCOMM 2023 · 被引用 63 次
相关 Paper
- From ATOP to ZCube: Automated Topology Optimization Pipeline and A Highly Cost-Effective Network Topology for Large Model TrainingZihan Yan, Dan Li, Li Chen, Dian Xiong 等SIGCOMM 2025 · 被引用 4 次
- PhOrch: Proactive Phase-Level Flow Path Orchestration For Contention-Free LLM TrainingZiyang Zou, Shuangwu Chen, Tao Zhang, Huihuang Qin 等INFOCOM 2026
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUsGuoliang He, Youhe Jiang, Wencong Xiao, Kaihua Jiang 等NeurIPS 2025 · 被引用 10 次
- 3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM TrainingHuifeng Xing, Hao Wang, Yinfan Hu, Xin Ai 等INFOCOM 2026
- InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching TransceiversChenchen Shou, Guyue Liu, Hao Nie, Huaiyu Meng 等SIGCOMM 2025 · 被引用 9 次
