ATP: In-network Aggregation for Multi-tenant Learning
ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen, Wenfei Wu, Aditya Akella, Michael M. Swift
Abstract
Distributed deep neural network training (DT) systems are widely deployed in clusters where the network is shared across multiple tenants, i.e., multiple DT jobs. Each DT job computes and aggregates gradients. Recent advances in hardware accelerators have shifted the the performance bottleneck of training from computation to communication. To speed up DT jobs' communication, we propose ATP, a service for in-network aggregation aimed at modern multi-rack, multi-job DT settings. ATP uses emerging programmable switch hardware to support in-network aggregation at multiple rack switches in a cluster to speedup DT jobs. ATP performs decentralized, dynamic, best-effort aggregation, enables efficient and equitable sharing of limited switch resources across simultaneously running DT jobs, and gracefully accommodates heavy contention for switch resources. ATP outperforms existing systems accelerating training throughput by up to 38% -66% in a cluster shared by multiple DT jobs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de9bccfb-1e63-4abe-8e39-09fc4ef6fed0Cited by top-tier papers65
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini et al.SIGCOMM 2021 · 120 citations
- Taurus: a data plane architecture for per-packet MLTushar Swamy, Alexander Rucker, Muhammad Shahbaz, Ishan Gaur et al.ASPLOS 2022 · 94 citations
- Brain-on-Switch: Towards Advanced Intelligent Network Data Plane via NN-Driven Traffic Analysis at Line-SpeedJinzhu Yan, Haotian Xu, Zhuotao Liu, Qi Li et al.NSDI 2024 · 60 citations
- Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingJiamin Cao, Yu Guan, Kun Qian, Jiaqi Gao et al.SIGCOMM 2024 · 60 citations
- Poseidon: Efficient, Robust, and Practical Datacenter CC via Deployable INTWeitao Wang, Masoud Moshref, Yuliang Li, Gautam Kumar et al.NSDI 2023 · 58 citations
Builds on5
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICsJiaqi Gao, Ennan Zhai, Hongqiang Harry Liu, Rui Miao et al.SIGCOMM 2020 · 82 citations
- Harmonia: Near-Linear Scalability for Replicated Storage with In-Network Conflict DetectionHang Zhu, Zhihao Bai, Jialin Li, Ellis Michael et al.VLDB 2020 · 58 citations
- Herring: rethinking the parameter server at scale for the cloudIndu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai et al.SC 2020 · 11 citations
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson et al.NSDI 2021
Related papers
- Training Job Placement in Clusters with Statistical In-Network AggregationBohan Zhao, Wei Xu, Shuo Liu, Yang Tian et al.ASPLOS 2024 · 17 citations
- InArt: In-Network Aggregation with Route Selection for Accelerating Distributed TrainingJiawei Liu, Yutong Zhai, Gongming Zhao, Hongli Xu et al.WWW 2024 · 13 citations
- Enabling Switch Memory Management for Distributed Training with In-Network AggregationBohan Zhao, Chang Liu, Jianbo Dong, Zheng Cao et al.INFOCOM 2023 · 22 citations
- Breaking Bucket Effect in In-Network Aggregation via Memory-Bandwidth CoordinationJunxu Xia, Geyao Cheng, Deke Guo, Lailong Luo et al.INFOCOM 2026
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 67 citations
