In-Network Aggregation with Transport Transparency for Distributed Training
Shuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu, Qinliang Lin, Yao Liu, Meng Xu, Marco Canini, Ray C. C. Cheung, Jianfei He
Abstract
Recent In-Network Aggregation (INA) solutions offload the allreduce operation onto network switches to accelerate and scale distributed training (DT). On end hosts, these solutions build custom network stacks to replace the transport layer. The INA-oriented network stack cannot take advantage of the state-of-the-art performant transport layer implementation, and also causes complexity in system development and operation.
We design a transport-transparent INA primitive named NetReduce for modern multi-rack data centers. NetReduce runs beneath the transport layer. The switch performs aggregation operations but preserves data transmission connections. The host uses RoCE as its transport layer to deliver gradient messages and receive aggregation results. NetReduce achieves performance gains from both INA and RoCE: linear scalability, traffic reduction, and bandwidth freeing-up from INA -high throughput, low latency, and low CPU overhead from RoCE. For jobs spanning several multi-GPU machines, we also devise parallel all-reduce based on NetReduce to make use of intra-machine and inter-machine bandwidth efficiently. We prototype NetReduce on an FPGA board attached to an Ethernet switch. We compare NetReduce with existing programmable switch-based solutions and justify the FPGA-based design choice. We evaluate NetReduce's performance by training typical Deep Neural Network models on single-GPU and multi-GPU testbeds. NetReduce inter-operates with the existing Ethernet transport layer,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 949fe10d-a55b-46a1-b05b-5e96f9b912afCited by top-tier papers10
- Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable MulticastWenxue Li, Junyi Zhang, Yufei Liu, Gaoxiong Zeng et al.HPCA 2024 · 13 citations
- MTP: Transport for In-Network ComputingTao Ji, Rohan Vardekar, Balajee Vamanan, Brent E. Stephens et al.NSDI 2025 · 9 citations
- SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA OffloadingXingyi Li, Yadong Liu, Xiaojie Huang, Yiran Zhang et al.NSDI 2026 · 2 citations
- A Generic and Efficient Communication Framework for Message-Level In-Network ComputingXinchen Wan, Luyang Li, Han Tian, Xudong Liao et al.INFOCOM 2025 · 2 citations
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li et al.ISCA 2026 · 1 citation
Builds on11
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee et al.OSDI 2020 · 286 citations
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini et al.SIGCOMM 2021 · 120 citations
- Elastic Resource Sharing for Distributed Deep LearningChangho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin et al.NSDI 2021 · 111 citations
- SiP-ML: high-bandwidth optical network interconnects for machine learning trainingMehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu et al.SIGCOMM 2021 · 94 citations
Related papers
- Training Job Placement in Clusters with Statistical In-Network AggregationBohan Zhao, Wei Xu, Shuo Liu, Yang Tian et al.ASPLOS 2024 · 17 citations
- Host-driven In-Network Aggregation on RDMAYulong Li, Wenxin Li, Yinan Yao, Yuxuan Du et al.INFOCOM 2024 · 1 citation
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 67 citations
- InArt: In-Network Aggregation with Route Selection for Accelerating Distributed TrainingJiawei Liu, Yutong Zhai, Gongming Zhao, Hongli Xu et al.WWW 2024 · 13 citations
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid et al.ISCA 2021 · 44 citations
