Asynchronous SGD Beats Minibatch SGD Under Arbitrary Delays
Konstantin Mishchenko, Francis R. Bach, Mathieu Even, Blake E. Woodworth
摘要
The existing analysis of asynchronous stochastic gradient descent (SGD) degrades dramatically when any delay is large, giving the impression that performance depends primarily on the delay. On the contrary, we prove much better guarantees for the same asynchronous SGD algorithm regardless of the delays in the gradients, depending instead just on the number of parallel devices used to implement the algorithm. Our guarantees are strictly better than the existing analyses, and we also argue that asynchronous SGD outperforms synchronous minibatch SGD in the settings we consider. For our analysis, we introduce a novel recursion based on "virtual iterates" and delay-adaptive stepsizes, which allow us to derive state-of-theart guarantees for both convex and non-convex objectives.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 被引用 131 次
- Stochastic Gradient Descent under Markovian Sampling SchemesMathieu EvenICML 2023 · 被引用 41 次
- Optimal Time Complexities of Parallel Stochastic Optimization Methods Under a Fixed Computation ModelAlexander Tyurin, Peter RichtárikNeurIPS 2023 · 被引用 31 次
- Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update CalibrationYujia Wang, Yuanpu Cao, Jingcheng Wu, Ruoyu Chen 等ICLR 2024 · 被引用 26 次
- Efficient Federated Learning against Heterogeneous and Non-stationary Client UnavailabilityMing Xiang, Stratis Ioannidis, Edmund Yeh, Carlee Joe-Wong 等NeurIPS 2024 · 被引用 26 次
它引用的顶会 Paper8
- FjORD: Fair and Accurate Federated Learning under heterogeneous targets with Ordered DropoutSamuel Horváth, Stefanos Laskaridis, Mário Almeida, Ilias Leontiadis 等NeurIPS 2021 · 被引用 390 次
- Is Local SGD Better than Minibatch SGD?Blake E. Woodworth, Kumar Kshitij Patel, Sebastian U. Stich, Zhen Dai 等ICML 2020 · 被引用 277 次
- FedAT: a high-performance and communication-efficient federated learning system with asynchronous tiersZheng Chai, Yujing Chen, Ali Anwar, Liang Zhao 等SC 2021 · 被引用 140 次
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 被引用 131 次
- Asynchronous Decentralized SGD with Quantized and Local UpdatesGiorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li 等NeurIPS 2021 · 被引用 61 次
相关 Paper
- Faster Stochastic Optimization with Arbitrary Delays via Adaptive Asynchronous Mini-BatchingAmit Attia, Ofir Gaash, Tomer KorenICML 2025
- Stability and Generalization of Asynchronous SGD: Sharper Bounds Beyond Lipschitz and SmoothnessXiaoge Deng, Tao Sun, Shengwei Li, Dongsheng Li 等NeurIPS 2024 · 被引用 3 次
- Delay-Adaptive Distributed Stochastic OptimizationZhaolin Ren, Zhengyuan Zhou, Linhai Qiu, Ajay Deshpande 等AAAI 2020 · 被引用 17 次
- Asynchronous Stochastic Optimization Robust to Arbitrary DelaysAlon Cohen, Amit Daniely, Yoel Drori, Tomer Koren 等NeurIPS 2021 · 被引用 46 次
- Delay-Adaptive Step-sizes for Asynchronous LearningXuyang Wu, Sindri Magnússon, Hamid Reza Feyzmahdavian, Mikael JohanssonICML 2022 · 被引用 17 次
