Prague: High-Performance Heterogeneity-Aware Asynchronous Decentralized Training
Qinyi Luo, Jiaao He, Youwei Zhuo, Xuehai Qian
Abstract
Distributed deep learning training usually adopts All-Reduce as the synchronization mechanism for data parallel algorithms due to its high performance in homogeneous environment. However, its performance is bounded by the slowest worker among all workers. For this reason, it is significantly slower in heterogeneous settings. AD-PSGD, a newly proposed synchronization method which provides numerically fast convergence and heterogeneity tolerance, suffers from deadlock issues and high synchronization overhead. Is it possible to get the best of both worlds --- designing a distributed training method that has both high performance like All-Reduce in homogeneous environment and good heterogeneity tolerance like AD-PSGD?
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a16f007f-eaff-45a3-8fe2-619b5c202690Cited by top-tier papers21
- IceBreaker: warming serverless functions better with heterogeneityRohan Basu Roy, Tirthak Patel, Devesh TiwariASPLOS 2022 · 151 citations
- SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural NetworksJingshu Peng, Zhao Chen, Yingxia Shao, Yanyan Shen et al.VLDB 2022 · 76 citations
- Communication-efficient Decentralized Machine Learning over Heterogeneous NetworksPan Zhou, Qian Lin, Dumitrel Loghin, Beng Chin Ooi et al.ICDE 2021 · 73 citations
- DecentLaM: Decentralized Momentum SGD for Large-batch Deep TrainingKun Yuan, Yiming Chen, Xinmeng Huang, Yingya Zhang et al.ICCV 2021 · 73 citations
- Adaptive Configuration for Heterogeneous Participants in Decentralized Federated LearningYunming Liao, Yang Xu, Hongli Xu, Lun Wang et al.INFOCOM 2023 · 66 citations
Related papers
- Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceXupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang et al.SIGMOD 2021 · 64 citations
- SDPipe: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel TrainingXupeng Miao, Yining Shi, Zhi Yang, Bin Cui et al.VLDB 2023 · 48 citations
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 14 citations
- Distributed Machine Learning through Heterogeneous Edge SystemsHanpeng Hu, Dan Wang, Chuan WuAAAI 2020 · 48 citations
- Addressing Network Bottlenecks with Divide-and-Shuffle Synchronization for Distributed DNN TrainingWeiyan Wang, Cengguang Zhang, Liu Yang, Kai Chen et al.INFOCOM 2022 · 14 citations
