Alibaba HPN: A Data Center Network for Large Language Model Training
Kun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao, Yichi Xu, Yu Guan, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao, Chao Wang, Peng Wang
Abstract
This paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. This requires us to design a new data center network architecture specifically for LLM training. Unlike general cloud computing which generates millions of small flows (e.g., lower than 10Gbps), LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP), the commonly used load-balancing scheme in traditional data centers, to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization by decreasing the occurrences of ECMP, but also greatly reduces the search space for path selection, thus allowing us to precisely select network paths capable of holding elephant flows. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to single-point failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks by addressing the layer-2 synchronization challenges. HPN has been deployed in our production for more than eight months. We share our experience in motivating, designing, and building HPN, as well as the operational lessons of HPN in production.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12e9839c-0c9c-487a-a54a-95ac14f0409cCited by top-tier papers47
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu et al.NSDI 2025 · 82 citations
- White-Boxing RDMA with Packet-Granular Software ControlChenxingyu Zhao, Jaehong Min, Ming Liu, Arvind KrishnamurthyNSDI 2025 · 28 citations
- Evolution of Aegis: Fault Diagnosis for AI Model Training Service in ProductionJianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng et al.NSDI 2025 · 21 citations
- Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep LearningWei An, Xiao Bi, Guanting Chen, Shanhuang Chen et al.SC 2024 · 20 citations
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang et al.USENIX ATC 2025 · 19 citations
Builds on6
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh et al.SIGCOMM 2022 · 230 citations
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre et al.NSDI 2023 · 117 citations
- PLB: congestion signals are simple and effective for network load balancingMubashir Adnan Qureshi, Yuchung Cheng, Qianwen Yin, Qiaobin Fu et al.SIGCOMM 2022 · 82 citations
- Lightwave Fabrics: At-Scale Optical Circuit Switching for Datacenter and Machine Learning SystemsHong Liu, Ryohei Urata, Kevin Yasumura, Xiang Zhou et al.SIGCOMM 2023 · 63 citations
Related papers
- From ATOP to ZCube: Automated Topology Optimization Pipeline and A Highly Cost-Effective Network Topology for Large Model TrainingZihan Yan, Dan Li, Li Chen, Dian Xiong et al.SIGCOMM 2025 · 4 citations
- PhOrch: Proactive Phase-Level Flow Path Orchestration For Contention-Free LLM TrainingZiyang Zou, Shuangwu Chen, Tao Zhang, Huihuang Qin et al.INFOCOM 2026
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUsGuoliang He, Youhe Jiang, Wencong Xiao, Kaihua Jiang et al.NeurIPS 2025 · 10 citations
- 3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM TrainingHuifeng Xing, Hao Wang, Yinfan Hu, Xin Ai et al.INFOCOM 2026
- InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching TransceiversChenchen Shou, Guyue Liu, Hao Nie, Huaiyu Meng et al.SIGCOMM 2025 · 9 citations
