Prague: High-Performance Heterogeneity-Aware Asynchronous Decentralized Training
Qinyi Luo, Jiaao He, Youwei Zhuo, Xuehai Qian
摘要
Distributed deep learning training usually adopts All-Reduce as the synchronization mechanism for data parallel algorithms due to its high performance in homogeneous environment. However, its performance is bounded by the slowest worker among all workers. For this reason, it is significantly slower in heterogeneous settings. AD-PSGD, a newly proposed synchronization method which provides numerically fast convergence and heterogeneity tolerance, suffers from deadlock issues and high synchronization overhead. Is it possible to get the best of both worlds --- designing a distributed training method that has both high performance like All-Reduce in homogeneous environment and good heterogeneity tolerance like AD-PSGD?
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper21
- IceBreaker: warming serverless functions better with heterogeneityRohan Basu Roy, Tirthak Patel, Devesh TiwariASPLOS 2022 · 被引用 151 次
- SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural NetworksJingshu Peng, Zhao Chen, Yingxia Shao, Yanyan Shen 等VLDB 2022 · 被引用 76 次
- Communication-efficient Decentralized Machine Learning over Heterogeneous NetworksPan Zhou, Qian Lin, Dumitrel Loghin, Beng Chin Ooi 等ICDE 2021 · 被引用 73 次
- DecentLaM: Decentralized Momentum SGD for Large-batch Deep TrainingKun Yuan, Yiming Chen, Xinmeng Huang, Yingya Zhang 等ICCV 2021 · 被引用 73 次
- Adaptive Configuration for Heterogeneous Participants in Decentralized Federated LearningYunming Liao, Yang Xu, Hongli Xu, Lun Wang 等INFOCOM 2023 · 被引用 66 次
相关 Paper
- Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceXupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang 等SIGMOD 2021 · 被引用 64 次
- SDPipe: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel TrainingXupeng Miao, Yining Shi, Zhi Yang, Bin Cui 等VLDB 2023 · 被引用 48 次
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 被引用 14 次
- Distributed Machine Learning through Heterogeneous Edge SystemsHanpeng Hu, Dan Wang, Chuan WuAAAI 2020 · 被引用 48 次
- Addressing Network Bottlenecks with Divide-and-Shuffle Synchronization for Distributed DNN TrainingWeiyan Wang, Cengguang Zhang, Liu Yang, Kai Chen 等INFOCOM 2022 · 被引用 14 次
