InArt: In-Network Aggregation with Route Selection for Accelerating Distributed Training
Jiawei Liu, Yutong Zhai, Gongming Zhao, Hongli Xu, Jin Fang, Zhen Zeng, Ying Zhu
Abstract
Deep learning has brought about a revolutionary transformation in network applications, particularly in domains like e-commerce and online advertising. Distributed training (DT), as a critical means to expedite model training, has progressively emerged as a key foundational infrastructure for such applications. However, with the rapid advancement of hardware accelerators, the performance bottleneck in DT has shifted from computation to communication. In-network aggregation (INA) solutions have shown promise in alleviating the communication bottleneck. Regrettably, current INA solutions primarily focus on improving efficiency under the traditional parameter server (PS) architecture and do not fully address the communication bottleneck caused by limited PS ingress bandwidth. To bridge this gap, we propose InArt, the first work to introduce INA with routing selection in a multi-PS architecture. InArt employs a multi-PS architecture to split DT tasks among multiple PSs, and selects appropriate routing schemes to fully harness INA capabilities. To accommodate traffic dynamics, InArt adopts a two-phase approach: splitting the training model among multiple parameter servers and selecting routing paths for INA. We propose Lagrange multiplier and randomized rounding algorithms for these phases, respectively. We implement InArt and evaluate its performance through experiments on physical platforms (Tofino switches) and Mininet emulation (P4 Software Switches). Experimental results show that InArt can reduce communication time by 48%!57!% compared with state-of-the-art solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c927636-bfca-40b2-b679-698711e68614Cited by top-tier papers2
- HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge NetworkPeirong Zheng, Wenchao Xu, Haozhao Wang, Jinyu Chen et al.INFOCOM 2026 · 2 citations
- Hierarchical Reinforcement Learning with Topology-Aware Exploration Framework for Multi-path Commodity Flow ProblemJingchen Jiang, Xuan Zhou, Jiayuan Li, Geng Han et al.AAAI 2026
Builds on3
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed TrainingChia-Yu Chen, Jiamin Ni, Songtao Lu, Xiaodong Cui et al.NeurIPS 2020 · 81 citations
Related papers
- A2TP: Aggregator-aware In-network Aggregation for Multi-tenant LearningZhaoyi Li, Jiawei Huang, Yijun Li, Aikun Xu et al.EuroSys 2023 · 35 citations
- 3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM TrainingHuifeng Xing, Hao Wang, Yinfan Hu, Xin Ai et al.INFOCOM 2026
- In-Network Aggregation with Transport Transparency for Distributed TrainingShuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu et al.ASPLOS 2023 · 46 citations
- Breaking Bucket Effect in In-Network Aggregation via Memory-Bandwidth CoordinationJunxu Xia, Geyao Cheng, Deke Guo, Lailong Luo et al.INFOCOM 2026
- Host-driven In-Network Aggregation on RDMAYulong Li, Wenxin Li, Yinan Yao, Yuxuan Du et al.INFOCOM 2024 · 1 citation
