GBA: A Tuning-free Approach to Switch between Synchronous and Asynchronous Training for Recommendation Models
Wenbo Su, Yuanxing Zhang, Yufeng Cai, Kaixu Ren, Pengjie Wang, Huimin Yi, Yue Song, Jing Chen, Hongbo Deng, Jian Xu, Lin Qu, Bo Zheng
Abstract
High-concurrency asynchronous training upon parameter server (PS) architecture and high-performance synchronous training upon all-reduce (AR) architecture are the most commonly deployed distributed training modes for recommendation models. Although synchronous AR training is designed to have higher training efficiency, asynchronous PS training would be a better choice for training speed when there are stragglers (slow workers) in the shared cluster, especially under limited computing resources. An ideal way to take full advantage of these two training modes is to switch between them upon the cluster status. However, switching training modes often requires tuning hyper-parameters, which is extremely time- and resource-consuming. We find two obstacles to a tuning-free approach: the different distribution of the gradient values and the stale gradients from the stragglers. This paper proposes Global Batch gradients Aggregation (GBA) over PS, which aggregates and applies gradients with the same global batch size as the synchronous training. A token-control process is implemented to assemble the gradients and decay the gradients with severe staleness. We provide the convergence analysis to reveal that GBA has comparable convergence properties with the synchronous training, and demonstrate the robustness of GBA the recommendation models against the gradient staleness. Experiments on three industrial-scale recommendation tasks show that GBA is an effective tuning-free approach for switching. Compared to the state-of-the-art derived asynchronous training, GBA achieves up to 0.2% improvement on the AUC metric, which is significant for the recommendation models. Meanwhile, under the strained hardware resource, GBA speeds up at least 2.4x compared to synchronous training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 548d1215-0a69-4fb8-8697-cbd4bd4c6f38Cited by top-tier papers3
- CO2: Efficient Distributed Training with Full Communication-Computation OverlapWeigao Sun, Zhen Qin, Weixuan Sun, Shidi Li et al.ICLR 2024 · 17 citations
- Provably Convergent Federated Trilevel LearningYang Jiao, Kai Yang, Tiancheng Wu, Chengtao Jian et al.AAAI 2024 · 6 citations
- Frugal: Efficient and Economic Embedding Model Training with Commodity GPUsMinhui Xie, Shaoxun Zeng, Hao Guo, Shiwei Gao et al.ASPLOS 2025 · 1 citation
Builds on9
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan et al.ICML 2020 · 125 citations
- Exponential Graph is Provably Efficient for Decentralized Deep TrainingBicheng Ying, Kun Yuan, Yiming Chen, Hanbin Hu et al.NeurIPS 2021 · 123 citations
- Improved Schemes for Episodic Memory-based Lifelong LearningYunhui Guo, Mingrui Liu, Tianbao Yang, Tajana RosingNeurIPS 2020 · 105 citations
Related papers
- Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningYijun Li, Jiawei Huang, Zhaoyi Li, Jingling Liu et al.INFOCOM 2024 · 2 citations
- Distributed Equivalent Substitution Training for Large-Scale Recommender SystemsHaidong Rong, Yangzihao Wang, Feihu Zhou, Junjie Zhai et al.SIGIR 2020 · 9 citations
- Learning Efficient Parameter Server Synchronization Policies for Distributed SGDRong Zhu, Sheng Yang, Andreas Pfadler, Zhengping Qian et al.ICLR 2020 · 9 citations
- Asynchronous Stochastic Gradient Descent for Extreme-Scale Recommender SystemsLewis Liu, Kun ZhaoAAAI 2021
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 14 citations
