Boosting Asynchronous Decentralized Learning with Model Fragmentation
Sayan Biswas, Anne-Marie Kermarrec, Alexis Marouani, Rafael Pires, Rishi Sharma, Martijn de Vos
Abstract
Decentralized learning (DL) is an emerging technique that allows nodes on the web to collaboratively train machine learning models without sharing raw data. Dealing with stragglers, i.e., nodes with slower compute or communication than others, is a key challenge in DL. We present DivShare, a novel asynchronous DL algorithm that achieves fast model convergence in the presence of communication stragglers. DivShare achieves this by having nodes fragment their models into parameter subsets and send, in parallel to computation, each subset to a random sample of other nodes instead of sequentially exchanging full models. The transfer of smaller fragments allows more efficient usage of the collective bandwidth and enables nodes with slow network links to quickly contribute with at least some of their model parameters. By theoretically proving the convergence of DivShare, we provide, to the best of our knowledge, the first formal proof of convergence for a DL algorithm that accounts for the effects of asynchronous communication with delays. We experimentally evaluate DivShare against two state-of-the-art DL baselines, AD-PSGD and Swift, and with two standard datasets, CIFAR-10 and MovieLens. We find that DivShare with communication stragglers lowers time-to-accuracy by up to 3.9× compared to AD-PSGD on the CIFAR-10 dataset. Compared to baselines, Di-vShare also achieves up to 19.4% better accuracy and 9.5% lower test loss on the CIFAR-10 and MovieLens datasets, respectively. CCS Concepts • Computing methodologies → Distributed artificial intelligence; Machine learning; Distributed algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e37d296-a170-4282-b72f-dd315f4f7f6aCited by top-tier papers1
Ask how each one uses itBuilds on16
- The Non-IID Data Quagmire of Decentralized Machine LearningKevin Hsieh, Amar Phanishayee, Onur Mutlu, Phillip B. GibbonsICML 2020 · 672 citations
- FedScale: Benchmarking Model and System Performance of Federated Learning at ScaleFan Lai, Yinwei Dai, Sanjay Sri Vallabh Singapuram, Jiachen Liu et al.ICML 2022 · 280 citations
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 263 citations
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 131 citations
- REFL: Resource-Efficient Federated LearningAhmed M. Abdelmoniem, Atal Narayan Sahu, Marco Canini, Suhaib A. FahmyEuroSys 2023 · 86 citations
Related papers
- SWIFT: Rapid Decentralized Federated Learning via Wait-Free Model CommunicationMarco Bornstein, Tahseen Rabbani, Evan Z. Wang, Amrit S. Bedi et al.ICLR 2023 · 3 citations
- Communication-efficient Decentralized Machine Learning over Heterogeneous NetworksPan Zhou, Qian Lin, Dumitrel Loghin, Beng Chin Ooi et al.ICDE 2021 · 73 citations
- Asynchronous Decentralized Online LearningJiyan Jiang, Wenpeng Zhang, Jinjie Gu, Wenwu ZhuNeurIPS 2021 · 26 citations
- Live Gradient Compensation for Evading Stragglers in Distributed LearningJian Xu, Shao-Lun Huang, Linqi Song, Tian LanINFOCOM 2021 · 29 citations
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 95 citations
