An Improved Analysis of Stochastic Gradient Descent with Momentum
Yanli Liu, Yuan Gao, Wotao Yin
Abstract
SGD with momentum (SGDM) has been widely applied in many machine learning tasks, and it is often applied with dynamic stepsizes and momentum weights tuned in a stagewise manner. Despite of its empirical advantage over SGD, the role of momentum is still unclear in general since previous analyses on SGDM either provide worse convergence bounds than those of SGD, or assume Lipschitz or quadratic objectives, which fail to hold in practice. Furthermore, the role of dynamic parameters have not been addressed. In this work, we show that SGDM converges as fast as SGD for smooth objectives under both strongly convex and nonconvex settings. We also establish the first convergence guarantee for the multistage setting, and show that the multistage strategy is beneficial for SGDM compared to using fixed parameters. Finally, we verify these theoretical claims by numerical experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49329a02-1ce2-4097-90b5-41ff81e9f0e8Cited by top-tier papers54
- Stochastic Controlled Averaging for Federated Learning with Communication CompressionXinmeng Huang, Ping Li, Xiaoyun LiICLR 2024 · 288 citations
- Learning from History for Byzantine Robust OptimizationSai Praneeth Karimireddy, Lie He, Martin JaggiICML 2021 · 247 citations
- Adam Can Converge Without Any Modification On Update RulesYushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun et al.NeurIPS 2022 · 134 citations
- High Probability Bounds for Non-Convex Stochastic Optimization with MomentumShaojie Li, Pengwei Tang, Bowei Zhu, Yong LiuICLR 2026 · 100 citations
- DecentLaM: Decentralized Momentum SGD for Large-batch Deep TrainingKun Yuan, Yiming Chen, Xinmeng Huang, Yingya Zhang et al.ICCV 2021 · 73 citations
Related papers
- On the Convergence of mSGD and AdaGrad for Stochastic OptimizationRuinan Jin, Yu Xing, Xingkang HeICLR 2022 · 12 citations
- Convergence of Distributed Adaptive Optimization with Local UpdatesZiheng Cheng, Margalit GlasgowICLR 2025
- Revisit last-iterate convergence of mSGD under milder requirement on step sizeRuinan Jin, Xingkang He, Lang Chen, Difei Cheng et al.NeurIPS 2022 · 6 citations
- Demystify Hyperparameters for Stochastic Optimization with Transferable RepresentationsJianhui Sun, Mengdi Huai, Kishlay Jha, Aidong ZhangKDD 2022 · 5 citations
- Adaptive Momentum by Momentum for Deep Neural Network TrainingTao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li et al.KDD 2026 · 1 citation
