Momentum Improves Normalized SGD
Ashok Cutkosky, Harsh Mehta
Abstract
We provide an improved analysis of normalized SGD showing that adding momentum provably removes the need for large batch sizes on non-convex objectives. Then, we consider the case of objectives with bounded second derivative and show that in this case a small tweak to the momentum formula allows normalized SGD with momentum to find an -critical point in iterations, matching the best-known rates without accruing any logarithmic factors or dependence on dimension. We also provide an adaptive method that automatically improves convergence rates when the variance in the gradients is small. Finally, we show that our method is effective when employed on popular large scale tasks such as ResNet-50 and BERT pretraining, matching the performance of the disparate methods used to get state-of-the-art results on both tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b4965c2-38d1-47ae-9f2c-5032795337b7Cited by top-tier papers69
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- Automatic Clipping: Differentially Private Deep Learning Made Easier and StrongerZhiqi Bu, Yu-Xiang Wang, Sheng Zha, George KarypisNeurIPS 2023 · 140 citations
- Improved Analysis of Clipping Algorithms for Non-convex OptimizationBohang Zhang, Jikai Jin, Cong Fang, Liwei WangNeurIPS 2020 · 139 citations
- High-probability Bounds for Non-Convex Stochastic Optimization with Heavy TailsAshok Cutkosky, Harsh MehtaNeurIPS 2021 · 119 citations
- Breaking the centralized barrier for cross-device federated learningSai Praneeth Karimireddy, Martin Jaggi, Satyen Kale, Mehryar Mohri et al.NeurIPS 2021 · 113 citations
Related papers
- Better SGD using Second-order MomentumHoang Tran, Ashok CutkoskyNeurIPS 2022 · 18 citations
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 25 citations
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han et al.ICLR 2021 · 165 citations
- ADOPT: Modified Adam Can Converge with Any β2 with the Optimal RateShohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima et al.NeurIPS 2024 · 32 citations
- Revisit last-iterate convergence of mSGD under milder requirement on step sizeRuinan Jin, Xingkang He, Lang Chen, Difei Cheng et al.NeurIPS 2022 · 6 citations
