Lune

INFOCOM2026顶会

3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM Training

Huifeng Xing, Hao Wang, Yinfan Hu, Xin Ai, Zixuan Chen, Yang Chen, Wanxin Shi, Sen Liu, Yang Xu

2026年份

摘要

Training large language models (LLMs) typically involves a technique known as 3D parallelism, which combines Data Parallelism, Tensor Parallelism, and Pipeline Parallelism. While effective, this method necessitates frequent, simultaneous all-reduce communications across multiple dimensions among training nodes, leading to substantial network traffic and congestion. To mitigate these issues, this paper explores the integration of In-network Aggregation (INA) into 3D parallelism. This integration encounters challenges due to conflicts in multi-dimensional communication patterns, topology-induced band-width underutilization, and the need for dynamic adaptation. In response, we propose a novel network architecture, 3D-INA, which directly incorporates traffic patterns into hardware and replaces traditional all-reduce communications with INA. This approach effectively bridges the gap between INA and the traffic patterns of 3D parallelism, while also resolving mismatches in network topology. Additionally, we employ reconfigurable optical switches and adaptive placement strategies, enhancing the system’s adaptation to varying workloads. Our simulations on the ns-3 platform demonstrate that 3D-INA can increase training throughput by up to six times compared to traditional topologies and all-reduce algorithms, while reducing costs and increasing flexibility. Further evaluations on a P4 testbed confirm the practical feasibility of 3D-INA for real-world applications, offering a promising new architecture for efficient LLM training.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get deb6fb96-0047-44b6-87a6-23cc7a33dc2a

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖