InArt: In-Network Aggregation with Route Selection for Accelerating Distributed Training
Jiawei Liu, Yutong Zhai, Gongming Zhao, Hongli Xu, Jin Fang, Zhen Zeng, Ying Zhu
摘要
Deep learning has brought about a revolutionary transformation in network applications, particularly in domains like e-commerce and online advertising. Distributed training (DT), as a critical means to expedite model training, has progressively emerged as a key foundational infrastructure for such applications. However, with the rapid advancement of hardware accelerators, the performance bottleneck in DT has shifted from computation to communication. In-network aggregation (INA) solutions have shown promise in alleviating the communication bottleneck. Regrettably, current INA solutions primarily focus on improving efficiency under the traditional parameter server (PS) architecture and do not fully address the communication bottleneck caused by limited PS ingress bandwidth. To bridge this gap, we propose InArt, the first work to introduce INA with routing selection in a multi-PS architecture. InArt employs a multi-PS architecture to split DT tasks among multiple PSs, and selects appropriate routing schemes to fully harness INA capabilities. To accommodate traffic dynamics, InArt adopts a two-phase approach: splitting the training model among multiple parameter servers and selecting routing paths for INA. We propose Lagrange multiplier and randomized rounding algorithms for these phases, respectively. We implement InArt and evaluate its performance through experiments on physical platforms (Tofino switches) and Mininet emulation (P4 Software Switches). Experimental results show that InArt can reduce communication time by 48%!57!% compared with state-of-the-art solutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge NetworkPeirong Zheng, Wenchao Xu, Haozhao Wang, Jinyu Chen 等INFOCOM 2026 · 被引用 2 次
- Hierarchical Reinforcement Learning with Topology-Aware Exploration Framework for Multi-path Commodity Flow ProblemJingchen Jiang, Xuan Zhou, Jiayuan Li, Geng Han 等AAAI 2026
它引用的顶会 Paper3
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed TrainingChia-Yu Chen, Jiamin Ni, Songtao Lu, Xiaodong Cui 等NeurIPS 2020 · 被引用 81 次
相关 Paper
- A2TP: Aggregator-aware In-network Aggregation for Multi-tenant LearningZhaoyi Li, Jiawei Huang, Yijun Li, Aikun Xu 等EuroSys 2023 · 被引用 35 次
- 3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM TrainingHuifeng Xing, Hao Wang, Yinfan Hu, Xin Ai 等INFOCOM 2026
- In-Network Aggregation with Transport Transparency for Distributed TrainingShuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu 等ASPLOS 2023 · 被引用 46 次
- Breaking Bucket Effect in In-Network Aggregation via Memory-Bandwidth CoordinationJunxu Xia, Geyao Cheng, Deke Guo, Lailong Luo 等INFOCOM 2026
- Host-driven In-Network Aggregation on RDMAYulong Li, Wenxin Li, Yinan Yao, Yuxuan Du 等INFOCOM 2024 · 被引用 1 次
