Two Sides of One Coin: the Limits of Untuned SGD and the Power of Adaptive Methods
Junchi Yang, Xiang Li, Ilyas Fatkhullin, Niao He
摘要
The classical analysis of Stochastic Gradient Descent (SGD) with polynomially decaying stepsize relies on well-tuned depending on problem parameters such as Lipschitz smoothness constant, which is often unknown in practice. In this work, we prove that SGD with arbitrary , referred to as untuned SGD, still attains an order-optimal convergence rate in terms of gradient norm for minimizing smooth objectives. Unfortunately, it comes at the expense of a catastrophic exponential dependence on the smoothness constant, which we show is unavoidable for this scheme even in the noiseless setting. We then examine three families of adaptive methods Normalized SGD (NSGD), AMSGrad, and AdaGrad unveiling their power in preventing such exponential dependency in the absence of information about the smoothness parameter and boundedness of stochastic gradients. Our results provide theoretical justification for the advantage of adaptive methods over untuned SGD in alleviating the issue with large gradients.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Adaptive Variance Reduction for Stochastic Optimization under Weaker AssumptionsWei Jiang, Sifan Yang, Yibo Wang, Lijun ZhangNeurIPS 2024 · 被引用 11 次
- Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity LimitsAbdurakhmon Sadiev, Peter Richtárik, Ilyas FatkhullinNeurIPS 2025 · 被引用 4 次
- Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGradZijian LiuICML 2026 · 被引用 3 次
- Problem-Parameter-Free Decentralized Bilevel OptimizationZhiwei Zhai, Wenjing Yan, Ying Jun ZhangNeurIPS 2025 · 被引用 2 次
- Momentum-Driven Adaptivity: Towards Tuning-Free Asynchronous Federated LearningWenjing Yan, Xiangyu Zhong, Xiaolu Wang, Ying-Jun Angela ZhangICML 2025
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong 等NeurIPS 2020 · 被引用 309 次
相关 Paper
- SGD with AdaGrad Stepsizes: Full Adaptivity with High Probability to Unknown Parameters, Unbounded Gradients and Affine VarianceAmit Attia, Tomer KorenICML 2023 · 被引用 34 次
- Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient DescentSharan Vaswani, Benjamin Dubois-Taine, Reza BabanezhadICML 2022
- High Probability Convergence of Stochastic Gradient MethodsZijian Liu, Ta Duy Nguyen, Thien Hang Nguyen, Alina Ene 等ICML 2023 · 被引用 64 次
- Stochastic Weakly Convex Optimization beyond Lipschitz ContinuityWenzhi Gao, Qi DengICML 2024 · 被引用 6 次
- High Probability Bounds for a Class of Nonconvex Algorithms with AdaGrad StepsizeAli Kavis, Kfir Yehuda Levy, Volkan CevherICLR 2022 · 被引用 51 次
