On Convergence-Diagnostic based Step Sizes for Stochastic Gradient Descent
Scott Pesme, Aymeric Dieuleveut, Nicolas Flammarion
Abstract
Constant step-size Stochastic Gradient Descent exhibits two phases: a transient phase during which iterates make fast progress towards the optimum, followed by a stationary phase during which iterates oscillate around the optimal point. In this paper, we show that efficiently detecting this transition and appropriately decreasing the step size can lead to fast convergence rates. We analyse the classical statistical test proposed by Pflug (1983) , based on the inner product between consecutive stochastic gradients. Even in the simple case where the objective function is quadratic we show that this test cannot lead to an adequate convergence diagnostic. We then propose a novel and simple statistical procedure that accurately detects stationarity and we provide experimental results showing state-of-the-art performance on synthetic and real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8a5ad2b-8d92-4576-826c-438ab68832ccCited by top-tier papers4
- Robust, Accurate Stochastic Optimization for Variational InferenceAkash Kumar Dhaka, Alejandro Catalina, Michael Riis Andersen, Måns Magnusson et al.NeurIPS 2020 · 39 citations
- Stochastic Reweighted Gradient DescentAyoub El Hanchi, David A. Stephens, Chris J. MaddisonICML 2022 · 10 citations
- Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient DescentXiang Li, Qiaomin XieAAAI 2025 · 1 citation
- Almost Bayesian: Dynamics of SGD Through Singular Learning TheoryMax Hennick, Stijn De BaerdemackerICLR 2026
Related papers
- Directional Smoothness and Gradient Methods: Convergence and AdaptivityAaron Mishkin, Ahmed Khaled, Yuanhao Wang, Aaron Defazio et al.NeurIPS 2024 · 25 citations
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 52 citations
- PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex OptimizationZhize Li, Hongyan Bao, Xiangliang Zhang, Peter RichtárikICML 2021 · 164 citations
- Derivatives of Stochastic Gradient Descent in parametric optimizationFranck Iutzeler, Edouard Pauwels, Samuel VaiterNeurIPS 2024
- Three-stage Evolution and Fast Equilibrium for SGD with Non-degerate Critical PointsYi Wang, Zhiren WangICML 2022 · 4 citations
