Adam Can Converge Without Any Modification On Update Rules
Yushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun, Zhi-Quan Luo
Abstract
Ever since Reddi et al. (2018) pointed out the divergence issue of Adam, many new variants have been designed to obtain convergence. However, vanilla Adam remains exceptionally popular and it works well in practice. Why is there a gap between theory and practice? We point out there is a mismatch between the settings of theory and practice: Reddi et al. ( 2018 ) pick the problem after picking the hyperparameters of Adam, i.e., (β 1 , β 2 ); while practical applications often fix the problem first and then tune (β 1 , β 2 ). Due to this observation, we conjecture that the empirical convergence can be theoretically justified, only if we change the order of picking the problem and hyperparameter. In this work, we confirm this conjecture. We prove that, when the 2nd-order momentum parameter β 2 is large and 1st-order momentum parameter β 1 < √ β 2 < 1, Adam converges to the neighborhood of critical points. The size of the neighborhood is propositional to the variance of stochastic gradients. Under an extra condition (strong growth condition), Adam converges to critical points. It is worth mentioning that our results cover a wide range of hyperparameters: as β 2 increases, our convergence result can cover any β 1 ∈ [0, 1) including β 1 = 0.9, which is the default setting in deep learning libraries. To our knowledge, this is the first result showing that Adam can converge without any modification on its update rules. Further, our analysis does not require assumptions of bounded gradients or bounded 2nd-order momentum. When β 2 is small, we further point out a large region of (β 1 , β 2 ) combinations where Adam can diverge to infinity. Our divergence result considers the same setting (fixing the optimization problem ahead) as our convergence result, indicating that there is a phase transition from divergence to convergence when increasing β 2 . These positive and negative results provide suggestions on how to tune Adam hyperparameters: for instance, when Adam does not work well, we suggest tuning up β 2 and trying β 1 < √ β 2 . * Correspondence author 2 We formally re-state their results in Appendix D.2. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e785e5c5-effa-45f7-a1ca-8c5a29045764Cited by top-tier papers46
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li et al.NeurIPS 2024 · 149 citations
- Convergence of Adam Under Relaxed AssumptionsHaochuan Li, Alexander Rakhlin, Ali JadbabaieNeurIPS 2023 · 132 citations
- Revisiting Gradient Clipping: Stochastic bias and tight convergence guaranteesAnastasia Koloskova, Hadrien Hendrikx, Sebastian U. StichICML 2023 · 106 citations
- Resolving Discrepancies in Compute-Optimal Scaling of Language ModelsTomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt et al.NeurIPS 2024 · 94 citations
- Implicit Bias of AdamW: ℓ∞-Norm Constrained OptimizationShuo Xie, Zhiyuan LiICML 2024 · 46 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- An Improved Analysis of Stochastic Gradient Descent with MomentumYanli Liu, Yuan Gao, Wotao YinNeurIPS 2020 · 328 citations
Related papers
- RMSprop converges with proper hyper-parameterNaichen Shi, Dawei Li, Mingyi Hong, Ruoyu SunICLR 2021 · 79 citations
- ADOPT: Modified Adam Can Converge with Any β2 with the Optimal RateShohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima et al.NeurIPS 2024 · 32 citations
- Towards Understanding Adam Convergence on Highly Degenerate PolynomialsZhiwei Bai, Jiajie Zhao, Zhangchen Zhou, Zhi-Qin John Xu et al.ICML 2026 · 2 citations
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 25 citations
- Closing the gap between the upper bound and lower bound of Adam's iteration complexityBohan Wang, Jingwen Fu, Huishuai Zhang, Nanning Zheng et al.NeurIPS 2023 · 2 citations
