Communication-efficient Decentralized Machine Learning over Heterogeneous Networks
Pan Zhou, Qian Lin, Dumitrel Loghin, Beng Chin Ooi, Yuncheng Wu, Hongfang Yu
摘要
In the last few years, distributed machine learning has been usually executed over heterogeneous networks such as a local area network within a multi-tenant cluster or a wide area network connecting data centers and edge clusters. In these heterogeneous networks, the link speeds among worker nodes vary significantly, making it challenging for state-of-the-art machine learning approaches to perform efficient training. Both centralized and decentralized training approaches suffer from low-speed links. In this paper, we propose a decentralized approach, namely NetMax, that enables worker nodes to communicate via high-speed links and, thus, significantly speed up the training process. NetMax possesses the following novel features. First, it consists of a novel consensus algorithm that allows worker nodes to train model copies on their local dataset asynchronously and exchange information via peer-to-peer communication to synchronize their local copies, instead of a central master node (i.e., parameter server). Second, each worker node selects one peer randomly with a fine-tuned probability to exchange information per iteration. In particular, peers with high-speed links are selected with high probability. Third, the probabilities of selecting peers are designed to minimize the total convergence time. Moreover, we mathematically prove the convergence of NetMax. We evaluate NetMax on heterogeneous cluster networks and show that it achieves speedups of 3.7×, 3.4×, and 1.9× in comparison with the state-of-the-art decentralized training approaches Prague, Allreduce-SGD, and AD-PSGD, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Adaptive Configuration for Heterogeneous Participants in Decentralized Federated LearningYunming Liao, Yang Xu, Hongli Xu, Lun Wang 等INFOCOM 2023 · 被引用 66 次
- Enhancing Federated Learning with In-Cloud Unlabeled DataLun Wang, Yang Xu, Hongli Xu, Jianchun Liu 等ICDE 2022 · 被引用 22 次
- SparDL: Distributed Deep Learning Training with Efficient Sparse CommunicationMinjun Zhao, Yichen Yin, Yuren Mao, Qing Liu 等ICDE 2024 · 被引用 6 次
它引用的顶会 Paper4
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen 等VLDB 2020 · 被引用 259 次
- Prague: High-Performance Heterogeneity-Aware Asynchronous Decentralized TrainingQinyi Luo, Jiaao He, Youwei Zhuo, Xuehai QianASPLOS 2020 · 被引用 73 次
- TRACER: A Framework for Facilitating Accurate and Interpretable Analytics for High Stakes ApplicationsKaiping Zheng, Shaofeng Cai, Horng Ruey Chua, Wei Wang 等SIGMOD 2020 · 被引用 22 次
- DLion: Decentralized Distributed Deep Learning in Micro-CloudsRankyung Hong, Abhishek ChandraHPDC 2021 · 被引用 20 次
相关 Paper
- Enhancing Decentralized Federated Learning for Non-IID Data on Heterogeneous DevicesMin Chen, Yang Xu, Hongli Xu, Liusheng HuangICDE 2023 · 被引用 25 次
- Quantized Decentralized Stochastic Learning over Directed GraphsHossein Taheri, Aryan Mokhtari, Hamed Hassani, Ramtin PedarsaniICML 2020 · 被引用 59 次
- Asynchronous Decentralized SGD with Quantized and Local UpdatesGiorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li 等NeurIPS 2021 · 被引用 61 次
- Heterogeneity-Aware Distributed Machine Learning Training via Partial ReduceXupeng Miao, Xiaonan Nie, Yingxia Shao, Zhi Yang 等SIGMOD 2021 · 被引用 64 次
- Distributed Machine Learning through Heterogeneous Edge SystemsHanpeng Hu, Dan Wang, Chuan WuAAAI 2020 · 被引用 48 次
