Boosting Asynchronous Decentralized Learning with Model Fragmentation
Sayan Biswas, Anne-Marie Kermarrec, Alexis Marouani, Rafael Pires, Rishi Sharma, Martijn de Vos
摘要
Decentralized learning (DL) is an emerging technique that allows nodes on the web to collaboratively train machine learning models without sharing raw data. Dealing with stragglers, i.e., nodes with slower compute or communication than others, is a key challenge in DL. We present DivShare, a novel asynchronous DL algorithm that achieves fast model convergence in the presence of communication stragglers. DivShare achieves this by having nodes fragment their models into parameter subsets and send, in parallel to computation, each subset to a random sample of other nodes instead of sequentially exchanging full models. The transfer of smaller fragments allows more efficient usage of the collective bandwidth and enables nodes with slow network links to quickly contribute with at least some of their model parameters. By theoretically proving the convergence of DivShare, we provide, to the best of our knowledge, the first formal proof of convergence for a DL algorithm that accounts for the effects of asynchronous communication with delays. We experimentally evaluate DivShare against two state-of-the-art DL baselines, AD-PSGD and Swift, and with two standard datasets, CIFAR-10 and MovieLens. We find that DivShare with communication stragglers lowers time-to-accuracy by up to 3.9× compared to AD-PSGD on the CIFAR-10 dataset. Compared to baselines, Di-vShare also achieves up to 19.4% better accuracy and 9.5% lower test loss on the CIFAR-10 and MovieLens datasets, respectively. CCS Concepts • Computing methodologies → Distributed artificial intelligence; Machine learning; Distributed algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- The Non-IID Data Quagmire of Decentralized Machine LearningKevin Hsieh, Amar Phanishayee, Onur Mutlu, Phillip B. GibbonsICML 2020 · 被引用 672 次
- FedScale: Benchmarking Model and System Performance of Federated Learning at ScaleFan Lai, Yinwei Dai, Sanjay Sri Vallabh Singapuram, Jiachen Liu 等ICML 2022 · 被引用 280 次
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 被引用 263 次
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 被引用 131 次
- REFL: Resource-Efficient Federated LearningAhmed M. Abdelmoniem, Atal Narayan Sahu, Marco Canini, Suhaib A. FahmyEuroSys 2023 · 被引用 86 次
相关 Paper
- SWIFT: Rapid Decentralized Federated Learning via Wait-Free Model CommunicationMarco Bornstein, Tahseen Rabbani, Evan Z. Wang, Amrit S. Bedi 等ICLR 2023 · 被引用 3 次
- Communication-efficient Decentralized Machine Learning over Heterogeneous NetworksPan Zhou, Qian Lin, Dumitrel Loghin, Beng Chin Ooi 等ICDE 2021 · 被引用 73 次
- Asynchronous Decentralized Online LearningJiyan Jiang, Wenpeng Zhang, Jinjie Gu, Wenwu ZhuNeurIPS 2021 · 被引用 26 次
- Live Gradient Compensation for Evading Stragglers in Distributed LearningJian Xu, Shao-Lun Huang, Linqi Song, Tian LanINFOCOM 2021 · 被引用 29 次
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 被引用 95 次
