ATP: In-network Aggregation for Multi-tenant Learning
ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen, Wenfei Wu, Aditya Akella, Michael M. Swift
摘要
Distributed deep neural network training (DT) systems are widely deployed in clusters where the network is shared across multiple tenants, i.e., multiple DT jobs. Each DT job computes and aggregates gradients. Recent advances in hardware accelerators have shifted the the performance bottleneck of training from computation to communication. To speed up DT jobs' communication, we propose ATP, a service for in-network aggregation aimed at modern multi-rack, multi-job DT settings. ATP uses emerging programmable switch hardware to support in-network aggregation at multiple rack switches in a cluster to speedup DT jobs. ATP performs decentralized, dynamic, best-effort aggregation, enables efficient and equitable sharing of limited switch resources across simultaneously running DT jobs, and gracefully accommodates heavy contention for switch resources. ATP outperforms existing systems accelerating training throughput by up to 38% -66% in a cluster shared by multiple DT jobs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper65
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini 等SIGCOMM 2021 · 被引用 120 次
- Taurus: a data plane architecture for per-packet MLTushar Swamy, Alexander Rucker, Muhammad Shahbaz, Ishan Gaur 等ASPLOS 2022 · 被引用 94 次
- Brain-on-Switch: Towards Advanced Intelligent Network Data Plane via NN-Driven Traffic Analysis at Line-SpeedJinzhu Yan, Haotian Xu, Zhuotao Liu, Qi Li 等NSDI 2024 · 被引用 60 次
- Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingJiamin Cao, Yu Guan, Kun Qian, Jiaqi Gao 等SIGCOMM 2024 · 被引用 60 次
- Poseidon: Efficient, Robust, and Practical Datacenter CC via Deployable INTWeitao Wang, Masoud Moshref, Yuliang Li, Gautam Kumar 等NSDI 2023 · 被引用 58 次
它引用的顶会 Paper5
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICsJiaqi Gao, Ennan Zhai, Hongqiang Harry Liu, Rui Miao 等SIGCOMM 2020 · 被引用 82 次
- Harmonia: Near-Linear Scalability for Replicated Storage with In-Network Conflict DetectionHang Zhu, Zhihao Bai, Jialin Li, Ellis Michael 等VLDB 2020 · 被引用 58 次
- Herring: rethinking the parameter server at scale for the cloudIndu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai 等SC 2020 · 被引用 11 次
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson 等NSDI 2021
相关 Paper
- Training Job Placement in Clusters with Statistical In-Network AggregationBohan Zhao, Wei Xu, Shuo Liu, Yang Tian 等ASPLOS 2024 · 被引用 17 次
- InArt: In-Network Aggregation with Route Selection for Accelerating Distributed TrainingJiawei Liu, Yutong Zhai, Gongming Zhao, Hongli Xu 等WWW 2024 · 被引用 13 次
- Enabling Switch Memory Management for Distributed Training with In-Network AggregationBohan Zhao, Chang Liu, Jianbo Dong, Zheng Cao 等INFOCOM 2023 · 被引用 22 次
- Breaking Bucket Effect in In-Network Aggregation via Memory-Bandwidth CoordinationJunxu Xia, Geyao Cheng, Deke Guo, Lailong Luo 等INFOCOM 2026
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 被引用 67 次
