On the Convergence of Step Decay Step-Size for Stochastic Optimization
Xiaoyu Wang, Sindri Magnússon, Mikael Johansson
摘要
The convergence of stochastic gradient descent is highly dependent on the step-size, especially on non-convex problems such as neural network training. Step decay step-size schedules (constant and then cut) are widely used in practice because of their excellent convergence and generalization qualities, but their theoretical properties are not yet well understood. We provide the convergence results for step decay in the non-convex regime, ensuring that the gradient norm vanishes at an rate. We also provide the convergence guarantees for general (possibly non-smooth) convex problems, ensuring an convergence rate. Finally, in the strongly convex case, we establish an rate for smooth problems, which we also prove to be tight, and an rate without the smoothness assumption. We illustrate the practical efficiency of the step decay step-size in several large scale deep neural network training tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Two Sides of One Coin: the Limits of Untuned SGD and the Power of Adaptive MethodsJunchi Yang, Xiang Li, Ilyas Fatkhullin, Niao HeNeurIPS 2023 · 被引用 33 次
- Generalized Polyak Step Size for First Order Optimization with MomentumXiaoyu Wang, Mikael Johansson, Tong ZhangICML 2023 · 被引用 32 次
- ADOPT: Modified Adam Can Converge with Any β2 with the Optimal RateShohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima 等NeurIPS 2024 · 被引用 32 次
- End-to-End Learning for Stochastic Optimization: A Bayesian PerspectiveYves Rychener, Daniel Kuhn, Tobias SutterICML 2023 · 被引用 14 次
- Provably Scalable Black-Box Variational Inference with Structured Variational FamiliesJoohwan Ko, Kyurae Kim, Woochang Kim, Jacob R. GardnerICML 2024 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)GradientsDimitris Oikonomou, Nicolas LoizouICML 2026 · 被引用 4 次
- Revisit last-iterate convergence of mSGD under milder requirement on step sizeRuinan Jin, Xingkang He, Lang Chen, Difei Cheng 等NeurIPS 2022 · 被引用 6 次
- A Second look at Exponential and Cosine Step Sizes: Simplicity, Adaptivity, and PerformanceXiaoyu Li, Zhenxun Zhuang, Francesco OrabonaICML 2021 · 被引用 29 次
- Error Feedback under (L0, L1)-Smoothness: Normalization and MomentumSarit Khirirat, Abdurakhmon Sadiev, Artem Riabinin, Eduard Gorbunov 等NeurIPS 2025 · 被引用 10 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
