Distributed Machine Learning through Heterogeneous Edge Systems
Hanpeng Hu, Dan Wang, Chuan Wu
Abstract
Many emerging AI applications request distributed machine learning (ML) among edge systems (e.g., IoT devices and PCs at the edge of the Internet), where data cannot be uploaded to a central venue for model training, due to their large volumes and/or security/privacy concerns. Edge devices are intrinsically heterogeneous in computing capacity, posing significant challenges to parameter synchronization for parallel training with the parameter server (PS) architecture. This paper proposes ADSP, a parameter synchronization model for distributed machine learning (ML) with heterogeneous edge systems. Eliminating the significant waiting time occurring with existing parameter synchronization models, the core idea of ADSP is to let faster edge devices continue training, while committing their model updates at strategically decided intervals. We design algorithms that decide time points for each worker to commit its model update, and ensure not only global model convergence but also faster convergence. Our testbed implementation and experiments show that ADSP outperforms existing parameter synchronization models significantly in terms of ML model convergence time, scalability and adaptability to large heterogeneity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 011cdc57-3bb3-4c32-bf8a-76b913ba9bf3Cited by top-tier papers2
- EdgeMove: Pipelining Device-Edge Model Training for Mobile IntelligenceZeqian Dong, Qiang He, Feifei Chen, Hai Jin et al.WWW 2023 · 12 citations
- ROG: A High Performance and Robust Distributed Training System for Robotic IoTXiuxian Guan, Zekai Sun, Shengliang Deng, Xusheng Chen et al.MICRO 2022 · 2 citations
Related papers
- Adaptive Configuration for Heterogeneous Participants in Decentralized Federated LearningYunming Liao, Yang Xu, Hongli Xu, Lun Wang et al.INFOCOM 2023 · 66 citations
- Learning Efficient Parameter Server Synchronization Policies for Distributed SGDRong Zhu, Sheng Yang, Andreas Pfadler, Zhengping Qian et al.ICLR 2020 · 9 citations
- FedMP: Federated Learning through Adaptive Model Pruning in Heterogeneous Edge ComputingZhida Jiang, Yang Xu, Hongli Xu, Zhiyuan Wang et al.ICDE 2022 · 86 citations
- ParallelSFL: A Novel Split Federated Learning Framework Tackling Heterogeneity IssuesYunming Liao, Yang Xu, Hongli Xu, Zhiwei Yao et al.MobiCom 2024 · 27 citations
- Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningYijun Li, Jiawei Huang, Zhaoyi Li, Jingling Liu et al.INFOCOM 2024 · 2 citations
