Gap-Aware Mitigation of Gradient Staleness
Saar Barkai, Ido Hakimi, Assaf Schuster
Abstract
Cloud computing is becoming increasingly popular as a platform for distributed training of deep neural networks. Synchronous stochastic gradient descent (SSGD) suffers from substantial slowdowns due to stragglers if the environment is non-dedicated, as is common in cloud computing. Asynchronous SGD (ASGD) methods are immune to these slowdowns but are scarcely used due to gradient staleness, which encumbers the convergence process. Recent techniques have had limited success mitigating the gradient staleness when scaling up to many workers (computing nodes). In this paper we define the Gap as a measure of gradient staleness and propose Gap-Aware (GA), a novel asynchronous-distributed method that penalizes stale gradients linearly to the Gap and performs well even when scaling to large numbers of workers. Our evaluation on the CIFAR, ImageNet, and WikiText-103 datasets shows that GA outperforms the currently acceptable gradient penalization method, in final test accuracy. We also provide convergence rate proof for GA. Despite prior beliefs, we show that if GA is applied, momentum becomes beneficial in asynchronous environments, even when the number of workers scales up.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- FLuID: Mitigating Stragglers in Federated Learning using Invariant DropoutIrene Wang, Prashant J. Nair, Divya MahajanNeurIPS 2023 · 42 citations
- SAPipe: Staleness-Aware Pipeline for Data Parallel DNN TrainingYangrui Chen, Cong Xie, Meng Ma, Juncheng Gu et al.NeurIPS 2022 · 24 citations
- Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelismSaar Eliad, Ido Hakimi, Alon De Jagger, Mark Silberstein et al.USENIX ATC 2021 · 24 citations
- CO2: Efficient Distributed Training with Full Communication-Computation OverlapWeigao Sun, Zhen Qin, Weixuan Sun, Shidi Li et al.ICLR 2024 · 17 citations
- MSPipe: Efficient Temporal GNN Training via Staleness-Aware PipelineGuangming Sheng, Junwei Su, Chao Huang, Chuan WuKDD 2024 · 7 citations
Related papers
- Ordered Momentum for Asynchronous SGDChang-Wei Shi, Yi-Rui Yang, Wu-Jun LiNeurIPS 2024 · 7 citations
- A2CiD2: Accelerating Asynchronous Communication in Decentralized Deep LearningAdel Nabli, Eugene Belilovsky, Edouard OyallonNeurIPS 2023 · 12 citations
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, Daniel SoudryICLR 2020 · 25 citations
- Asynchronous Optimization Methods for Efficient Training of Deep Neural Networks with GuaranteesVyacheslav Kungurtsev, Malcolm Egan, Bapi Chatterjee, Dan AlistarhAAAI 2021 · 4 citations
- Ordered Local Momentum for Asynchronous Distributed Learning Under Arbitrary DelaysChang-Wei Shi, Shi-Shang Wang, Wu-Jun LiAAAI 2026
