Asynchronous SGD Beats Minibatch SGD Under Arbitrary Delays
Konstantin Mishchenko, Francis R. Bach, Mathieu Even, Blake E. Woodworth
Abstract
The existing analysis of asynchronous stochastic gradient descent (SGD) degrades dramatically when any delay is large, giving the impression that performance depends primarily on the delay. On the contrary, we prove much better guarantees for the same asynchronous SGD algorithm regardless of the delays in the gradients, depending instead just on the number of parallel devices used to implement the algorithm. Our guarantees are strictly better than the existing analyses, and we also argue that asynchronous SGD outperforms synchronous minibatch SGD in the settings we consider. For our analysis, we introduce a novel recursion based on "virtual iterates" and delay-adaptive stepsizes, which allow us to derive state-of-theart guarantees for both convex and non-convex objectives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81bc8e7c-ddd5-450f-a1e4-aff8f33a4a43Cited by top-tier papers32
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 131 citations
- Stochastic Gradient Descent under Markovian Sampling SchemesMathieu EvenICML 2023 · 41 citations
- Optimal Time Complexities of Parallel Stochastic Optimization Methods Under a Fixed Computation ModelAlexander Tyurin, Peter RichtárikNeurIPS 2023 · 31 citations
- Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update CalibrationYujia Wang, Yuanpu Cao, Jingcheng Wu, Ruoyu Chen et al.ICLR 2024 · 26 citations
- Efficient Federated Learning against Heterogeneous and Non-stationary Client UnavailabilityMing Xiang, Stratis Ioannidis, Edmund Yeh, Carlee Joe-Wong et al.NeurIPS 2024 · 26 citations
Builds on8
- FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered DropoutSamuel Horváth, Stefanos Laskaridis, Mário Almeida, Ilias Leontiadis et al.NeurIPS 2021 · 390 citations
- Is Local SGD Better than Minibatch SGD?Blake E. Woodworth, Kumar Kshitij Patel, Sebastian U. Stich, Zhen Dai et al.ICML 2020 · 277 citations
- FedAT: a high-performance and communication-efficient federated learning system with asynchronous tiersZheng Chai, Yujing Chen, Ali Anwar, Liang Zhao et al.SC 2021 · 140 citations
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 131 citations
- Asynchronous Decentralized SGD with Quantized and Local UpdatesGiorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li et al.NeurIPS 2021 · 61 citations
Related papers
- Faster Stochastic Optimization with Arbitrary Delays via Adaptive Asynchronous Mini-BatchingAmit Attia, Ofir Gaash, Tomer KorenICML 2025
- Stability and Generalization of Asynchronous SGD: Sharper Bounds Beyond Lipschitz and SmoothnessXiaoge Deng, Tao Sun, Shengwei Li, Dongsheng Li et al.NeurIPS 2024 · 3 citations
- Delay-Adaptive Distributed Stochastic OptimizationZhaolin Ren, Zhengyuan Zhou, Linhai Qiu, Ajay Deshpande et al.AAAI 2020 · 17 citations
- Asynchronous Stochastic Optimization Robust to Arbitrary DelaysAlon Cohen, Amit Daniely, Yoel Drori, Tomer Koren et al.NeurIPS 2021 · 46 citations
- Delay-Adaptive Step-sizes for Asynchronous LearningXuyang Wu, Sindri Magnússon, Hamid Reza Feyzmahdavian, Mikael JohanssonICML 2022 · 17 citations
