Communication-efficient Decentralized Machine Learning over Heterogeneous Networks
Pan Zhou, Qian Lin, Dumitrel Loghin, Beng Chin Ooi, Yuncheng Wu, Hongfang Yu
Abstract
In the last few years, distributed machine learning has been usually executed over heterogeneous networks such as a local area network within a multi-tenant cluster or a wide area network connecting data centers and edge clusters. In these heterogeneous networks, the link speeds among worker nodes vary significantly, making it challenging for state-of-the-art machine learning approaches to perform efficient training. Both centralized and decentralized training approaches suffer from low-speed links. In this paper, we propose a decentralized approach, namely NetMax, that enables worker nodes to communicate via high-speed links and, thus, significantly speed up the training process. NetMax possesses the following novel features. First, it consists of a novel consensus algorithm that allows worker nodes to train model copies on their local dataset asynchronously and exchange information via peer-to-peer communication to synchronize their local copies, instead of a central master node (i.e., parameter server). Second, each worker node selects one peer randomly with a fine-tuned probability to exchange information per iteration. In particular, peers with high-speed links are selected with high probability. Third, the probabilities of selecting peers are designed to minimize the total convergence time. Moreover, we mathematically prove the convergence of NetMax. We evaluate NetMax on heterogeneous cluster networks and show that it achieves speedups of 3.7×, 3.4×, and 1.9× in comparison with the state-of-the-art decentralized training approaches Prague, Allreduce-SGD, and AD-PSGD, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6552b7d9-1392-4318-85b4-e94332ca2a16Cited by top-tier papers3
- Adaptive Configuration for Heterogeneous Participants in Decentralized Federated LearningYunming Liao, Yang Xu, Hongli Xu, Lun Wang et al.INFOCOM 2023 · 66 citations
- Enhancing Federated Learning with In-Cloud Unlabeled DataLun Wang, Yang Xu, Hongli Xu, Jianchun Liu et al.ICDE 2022 · 22 citations
- SparDL: Distributed Deep Learning Training with Efficient Sparse CommunicationMinjun Zhao, Yichen Yin, Yuren Mao, Qing Liu et al.ICDE 2024 · 6 citations
Builds on4
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen et al.VLDB 2020 · 259 citations
- Prague: High-Performance Heterogeneity-Aware Asynchronous Decentralized TrainingQinyi Luo, Jiaao He, Youwei Zhuo, Xuehai QianASPLOS 2020 · 73 citations
- TRACER: A Framework for Facilitating Accurate and Interpretable Analytics for High Stakes ApplicationsKaiping Zheng, Shaofeng Cai, Horng Ruey Chua, Wei Wang et al.SIGMOD 2020 · 22 citations
- DLion: Decentralized Distributed Deep Learning in Micro-CloudsRankyung Hong, Abhishek ChandraHPDC 2021 · 20 citations
Related papers
- Enhancing Decentralized Federated Learning for Non-IID Data on Heterogeneous DevicesMin Chen, Yang Xu, Hongli Xu, Liusheng HuangICDE 2023 · 25 citations
- Quantized Decentralized Stochastic Learning over Directed GraphsHossein Taheri, Aryan Mokhtari, Hamed Hassani, Ramtin PedarsaniICML 2020 · 59 citations
- Asynchronous Decentralized SGD with Quantized and Local UpdatesGiorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li et al.NeurIPS 2021 · 61 citations
- Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceXupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang et al.SIGMOD 2021 · 64 citations
- Distributed Machine Learning through Heterogeneous Edge SystemsHanpeng Hu, Dan Wang, Chuan WuAAAI 2020 · 48 citations
