ASAP.SGD: Instance-based Adaptiveness to Staleness in Asynchronous SGD
Karl Bäckström, Marina Papatriantafilou, Philippas Tsigas
Abstract
Concurrent algorithmic implementations of Stochastic Gradient Descent (SGD) give rise to critical questions for compute-intensive Machine Learning (ML). Asynchrony implies speedup in some contexts, and challenges in others, as stale updates may lead to slower, or nonconverging executions. While previous works showed asynchrony-adaptiveness can improve stability and speedup by reducing the step size for stale updates according to static rules, there is no one-size-fits-all adaptation rule, since the optimal strategy depends on several factors. We introduce (i) ASAP.SGD, an analytical framework capturing necessary and desired properties of stalenessadaptive step size functions and (ii) TAIL-τ , a method for utilizing key properties of the execution instance, generating a tailored strategy that not only dampens the impact of stale updates, but also leverages fresh ones. We recover convergence bounds for adaptiveness functions satisfying the ASAP.SGD conditions, for general, convex and non-convex problems, and establish novel bounds for ones satisfying the Polyak-Lojasiewicz property. We evaluate TAIL-τ with representative AsyncSGD concurrent algorithms, for Deep Learning problems, showing TAIL-τ is a vital complement to AsyncSGD, with (i) persistent speedup in wall-clock convergence time in the parallelism spectrum, (ii) considerably lower risk of non-convergence, as well as (iii) precision levels for which original SGD implementations fail.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f393e264-a586-4b3f-8bbe-bbac0bbf5912Builds on2
- Asynchronous Distributed Learning : Adapting to Gradient Delays without Prior KnowledgeRotem Zamir Aviv, Ido Hakimi, Assaf Schuster, Kfir Yehuda LevyICML 2021 · 21 citations
- Delay-Adaptive Distributed Stochastic OptimizationZhaolin Ren, Zhengyuan Zhou, Linhai Qiu, Ajay Deshpande et al.AAAI 2020 · 17 citations
Related papers
- Gap-Aware Mitigation of Gradient StalenessSaar Barkai, Ido Hakimi, Assaf SchusterICLR 2020 · 27 citations
- Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)GradientsDimitris Oikonomou, Nicolas LoizouICML 2026 · 4 citations
- STL-SGD: Speeding Up Local SGD with Stagewise Communication PeriodShuheng Shen, Yifei Cheng, Jingchang Liu, Linli XuAAAI 2021 · 12 citations
- A Second look at Exponential and Cosine Step Sizes: Simplicity, Adaptivity, and PerformanceXiaoyu Li, Zhenxun Zhuang, Francesco OrabonaICML 2021 · 29 citations
- Robustness Analysis of Non-Convex Stochastic Gradient Descent using Biased ExpectationsKevin Scaman, Cédric MalherbeNeurIPS 2020 · 37 citations
