In-Network Aggregation with Transport Transparency for Distributed Training
Shuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu, Qinliang Lin, Yao Liu, Meng Xu, Marco Canini, Ray C. C. Cheung, Jianfei He
摘要
Recent In-Network Aggregation (INA) solutions offload the allreduce operation onto network switches to accelerate and scale distributed training (DT). On end hosts, these solutions build custom network stacks to replace the transport layer. The INA-oriented network stack cannot take advantage of the state-of-the-art performant transport layer implementation, and also causes complexity in system development and operation.
We design a transport-transparent INA primitive named NetReduce for modern multi-rack data centers. NetReduce runs beneath the transport layer. The switch performs aggregation operations but preserves data transmission connections. The host uses RoCE as its transport layer to deliver gradient messages and receive aggregation results. NetReduce achieves performance gains from both INA and RoCE: linear scalability, traffic reduction, and bandwidth freeing-up from INA -high throughput, low latency, and low CPU overhead from RoCE. For jobs spanning several multi-GPU machines, we also devise parallel all-reduce based on NetReduce to make use of intra-machine and inter-machine bandwidth efficiently. We prototype NetReduce on an FPGA board attached to an Ethernet switch. We compare NetReduce with existing programmable switch-based solutions and justify the FPGA-based design choice. We evaluate NetReduce's performance by training typical Deep Neural Network models on single-GPU and multi-GPU testbeds. NetReduce inter-operates with the existing Ethernet transport layer,
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable MulticastWenxue Li, Junyi Zhang, Yufei Liu, Gaoxiong Zeng 等HPCA 2024 · 被引用 13 次
- MTP: Transport for In-Network ComputingTao Ji, Rohan Vardekar, Balajee Vamanan, Brent E. Stephens 等NSDI 2025 · 被引用 9 次
- SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA OffloadingXingyi Li, Yadong Liu, Xiaojie Huang, Yiran Zhang 等NSDI 2026 · 被引用 2 次
- A Generic and Efficient Communication Framework for Message-Level In-Network ComputingXinchen Wan, Luyang Li, Han Tian, Xudong Liao 等INFOCOM 2025 · 被引用 2 次
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li 等ISCA 2026 · 被引用 1 次
它引用的顶会 Paper11
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee 等OSDI 2020 · 被引用 286 次
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini 等SIGCOMM 2021 · 被引用 120 次
- Elastic Resource Sharing for Distributed Deep LearningChangho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin 等NSDI 2021 · 被引用 111 次
- SiP-ML: high-bandwidth optical network interconnects for machine learning trainingMehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu 等SIGCOMM 2021 · 被引用 94 次
相关 Paper
- Training Job Placement in Clusters with Statistical In-Network AggregationBohan Zhao, Wei Xu, Shuo Liu, Yang Tian 等ASPLOS 2024 · 被引用 17 次
- Host-driven In-Network Aggregation on RDMAYulong Li, Wenxin Li, Yinan Yao, Yuxuan Du 等INFOCOM 2024 · 被引用 1 次
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 被引用 67 次
- InArt: In-Network Aggregation with Route Selection for Accelerating Distributed TrainingJiawei Liu, Yutong Zhai, Gongming Zhao, Hongli Xu 等WWW 2024 · 被引用 13 次
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid 等ISCA 2021 · 被引用 44 次
