GBA: A Tuning-free Approach to Switch between Synchronous and Asynchronous Training for Recommendation Models
Wenbo Su, Yuanxing Zhang, Yufeng Cai, Kaixu Ren, Pengjie Wang, Huimin Yi, Yue Song, Jing Chen, Hongbo Deng, Jian Xu, Lin Qu, Bo Zheng
摘要
High-concurrency asynchronous training upon parameter server (PS) architecture and high-performance synchronous training upon all-reduce (AR) architecture are the most commonly deployed distributed training modes for recommendation models. Although synchronous AR training is designed to have higher training efficiency, asynchronous PS training would be a better choice for training speed when there are stragglers (slow workers) in the shared cluster, especially under limited computing resources. An ideal way to take full advantage of these two training modes is to switch between them upon the cluster status. However, switching training modes often requires tuning hyper-parameters, which is extremely time- and resource-consuming. We find two obstacles to a tuning-free approach: the different distribution of the gradient values and the stale gradients from the stragglers. This paper proposes Global Batch gradients Aggregation (GBA) over PS, which aggregates and applies gradients with the same global batch size as the synchronous training. A token-control process is implemented to assemble the gradients and decay the gradients with severe staleness. We provide the convergence analysis to reveal that GBA has comparable convergence properties with the synchronous training, and demonstrate the robustness of GBA the recommendation models against the gradient staleness. Experiments on three industrial-scale recommendation tasks show that GBA is an effective tuning-free approach for switching. Compared to the state-of-the-art derived asynchronous training, GBA achieves up to 0.2% improvement on the AUC metric, which is significant for the recommendation models. Meanwhile, under the strained hardware resource, GBA speeds up at least 2.4x compared to synchronous training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CO2: Efficient Distributed Training with Full Communication-Computation OverlapWeigao Sun, Zhen Qin, Weixuan Sun, Shidi Li 等ICLR 2024 · 被引用 17 次
- Provably Convergent Federated Trilevel LearningYang Jiao, Kai Yang, Tiancheng Wu, Chengtao Jian 等AAAI 2024 · 被引用 6 次
- Frugal: Efficient and Economic Embedding Model Training with Commodity GPUsMinhui Xie, Shaoxun Zeng, Hao Guo, Shiwei Gao 等ASPLOS 2025 · 被引用 1 次
它引用的顶会 Paper9
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan 等ICML 2020 · 被引用 125 次
- Exponential Graph is Provably Efficient for Decentralized Deep TrainingBicheng Ying, Kun Yuan, Yiming Chen, Hanbin Hu 等NeurIPS 2021 · 被引用 123 次
- Improved Schemes for Episodic Memory-based Lifelong LearningYunhui Guo, Mingrui Liu, Tianbao Yang, Tajana RosingNeurIPS 2020 · 被引用 105 次
相关 Paper
- Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningYijun Li, Jiawei Huang, Zhaoyi Li, Jingling Liu 等INFOCOM 2024 · 被引用 2 次
- Distributed Equivalent Substitution Training for Large-Scale Recommender SystemsHaidong Rong, Yangzihao Wang, Feihu Zhou, Junjie Zhai 等SIGIR 2020 · 被引用 9 次
- Learning Efficient Parameter Server Synchronization Policies for Distributed SGDRong Zhu, Sheng Yang, Andreas Pfadler, Zhengping Qian 等ICLR 2020 · 被引用 9 次
- Asynchronous Stochastic Gradient Descent for Extreme-Scale Recommender SystemsLewis Liu, Kun ZhaoAAAI 2021
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 被引用 14 次
