3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM Training
Huifeng Xing, Hao Wang, Yinfan Hu, Xin Ai, Zixuan Chen, Yang Chen, Wanxin Shi, Sen Liu, Yang Xu
摘要
Training large language models (LLMs) typically involves a technique known as 3D parallelism, which combines Data Parallelism, Tensor Parallelism, and Pipeline Parallelism. While effective, this method necessitates frequent, simultaneous all-reduce communications across multiple dimensions among training nodes, leading to substantial network traffic and congestion. To mitigate these issues, this paper explores the integration of In-network Aggregation (INA) into 3D parallelism. This integration encounters challenges due to conflicts in multi-dimensional communication patterns, topology-induced band-width underutilization, and the need for dynamic adaptation. In response, we propose a novel network architecture, 3D-INA, which directly incorporates traffic patterns into hardware and replaces traditional all-reduce communications with INA. This approach effectively bridges the gap between INA and the traffic patterns of 3D parallelism, while also resolving mismatches in network topology. Additionally, we employ reconfigurable optical switches and adaptive placement strategies, enhancing the system’s adaptation to varying workloads. Our simulations on the ns-3 platform demonstrate that 3D-INA can increase training throughput by up to six times compared to traditional topologies and all-reduce algorithms, while reducing costs and increasing flexibility. Further evaluations on a P4 testbed confirm the practical feasibility of 3D-INA for real-world applications, offering a promising new architecture for efficient LLM training.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- InArt: In-Network Aggregation with Route Selection for Accelerating Distributed TrainingJiawei Liu, Yutong Zhai, Gongming Zhao, Hongli Xu 等WWW 2024 · 被引用 13 次
- Host-driven In-Network Aggregation on RDMAYulong Li, Wenxin Li, Yinan Yao, Yuxuan Du 等INFOCOM 2024 · 被引用 1 次
- A2TP: Aggregator-aware In-network Aggregation for Multi-tenant LearningZhaoyi Li, Jiawei Huang, Yijun Li, Aikun Xu 等EuroSys 2023 · 被引用 35 次
- In-Network Aggregation with Transport Transparency for Distributed TrainingShuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu 等ASPLOS 2023 · 被引用 46 次
- WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model TrainingZheng Wang, Anna Cai, Xinfeng Xie, Zaifeng Pan 等OSDI 2025 · 被引用 22 次
