On Convergence-Diagnostic based Step Sizes for Stochastic Gradient Descent
Scott Pesme, Aymeric Dieuleveut, Nicolas Flammarion
摘要
Constant step-size Stochastic Gradient Descent exhibits two phases: a transient phase during which iterates make fast progress towards the optimum, followed by a stationary phase during which iterates oscillate around the optimal point. In this paper, we show that efficiently detecting this transition and appropriately decreasing the step size can lead to fast convergence rates. We analyse the classical statistical test proposed by Pflug (1983) , based on the inner product between consecutive stochastic gradients. Even in the simple case where the objective function is quadratic we show that this test cannot lead to an adequate convergence diagnostic. We then propose a novel and simple statistical procedure that accurately detects stationarity and we provide experimental results showing state-of-the-art performance on synthetic and real-world datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Robust, Accurate Stochastic Optimization for Variational InferenceAkash Kumar Dhaka, Alejandro Catalina, Michael Riis Andersen, Måns Magnusson 等NeurIPS 2020 · 被引用 39 次
- Stochastic Reweighted Gradient DescentAyoub El Hanchi, David A. Stephens, Chris J. MaddisonICML 2022 · 被引用 10 次
- Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient DescentXiang Li, Qiaomin XieAAAI 2025 · 被引用 1 次
- Almost Bayesian: Dynamics of SGD Through Singular Learning TheoryMax Hennick, Stijn De BaerdemackerICLR 2026
相关 Paper
- Directional Smoothness and Gradient Methods: Convergence and AdaptivityAaron Mishkin, Ahmed Khaled, Yuanhao Wang, Aaron Defazio 等NeurIPS 2024 · 被引用 25 次
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 52 次
- PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex OptimizationZhize Li, Hongyan Bao, Xiangliang Zhang, Peter RichtárikICML 2021 · 被引用 164 次
- Derivatives of Stochastic Gradient Descent in parametric optimizationFranck Iutzeler, Edouard Pauwels, Samuel VaiterNeurIPS 2024
- Three-stage Evolution and Fast Equilibrium for SGD with Non-degerate Critical PointsYi Wang, Zhiren WangICML 2022 · 被引用 4 次
