Learning Efficient Parameter Server Synchronization Policies for Distributed SGD
Rong Zhu, Sheng Yang, Andreas Pfadler, Zhengping Qian, Jingren Zhou
Abstract
We apply a reinforcement learning (RL) based approach to learning optimal synchronization policies used for Parameter Server-based distributed training of machine learning models with Stochastic Gradient Descent (SGD). Utilizing a formal synchronization policy description in the PS-setting, we are able to derive a suitable and compact description of states and actions, allowing us to efficiently use the standard off-the-shelf deep Q-learning algorithm. As a result, we are able to learn synchronization policies which generalize to different cluster environments, different training datasets and small model variations and (most importantly) lead to considerable decreases in training time when compared to standard policies such as bulk synchronous parallel (BSP), asynchronous parallel (ASP), or stale synchronous parallel (SSP). To support our claims we present extensive numerical results obtained from experiments performed in simulated cluster environments. In our experiments training time is reduced by 44 on average and learned policies generalize to multiple unseen circumstances.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 109b0fa6-aedd-4d16-bc55-b2fbef8f37e4Related papers
- Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningYijun Li, Jiawei Huang, Zhaoyi Li, Jingling Liu et al.INFOCOM 2024 · 2 citations
- Distributed Machine Learning through Heterogeneous Edge SystemsHanpeng Hu, Dan Wang, Chuan WuAAAI 2020 · 48 citations
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 14 citations
- AutoSync: Learning to Synchronize for Data-Parallel Distributed Deep LearningHao Zhang, Yuan Li, Zhijie Deng, Xiaodan Liang et al.NeurIPS 2020 · 33 citations
- GBA: A Tuning-free Approach to Switch between Synchronous and Asynchronous Training for Recommendation ModelsWenbo Su, Yuanxing Zhang, Yufeng Cai, Kaixu Ren et al.NeurIPS 2022 · 6 citations
