Convergence of Adam Under Relaxed Assumptions
Haochuan Li, Alexander Rakhlin, Ali Jadbabaie
Abstract
In this paper, we provide a rigorous proof of convergence of the Adaptive Moment Estimate (Adam) algorithm for a wide class of optimization objectives. Despite the popularity and efficiency of the Adam algorithm in training deep neural networks, its theoretical properties are not yet fully understood, and existing convergence proofs require unrealistically strong assumptions, such as globally bounded gradients, to show the convergence to stationary points. In this paper, we show that Adam provably converges to ǫ-stationary points with O(ǫ -4 ) gradient complexity under far more realistic conditions. The key to our analysis is a new proof of boundedness of gradients along the optimization trajectory of Adam, under a generalized smoothness assumption according to which the local smoothness (i.e., Hessian norm when it exists) is bounded by a sub-quadratic function of the gradient norm. Moreover, we propose a variance-reduced version of Adam with an accelerated gradient complexity of O(ǫ -3 ). not converge. That being said, the counter-examples depend on the hyper-parameters of Adam, i.e., they are constructed after picking the hyper-parameters. Therefore, it does not rule out the possibility of obtaining convergence guarantees for problem-dependent hyper-parameters, as also pointed out by (Shi et al., 2021; Zhang et al., 2022) . Many recent works have developed convergence analyses of Adam with various assumptions and hyperparameter choices. Zhou et al. (2018b) show Adam with certain hyper-parameters can work on the counterexamples of (Reddi et al., 2018) . De et al. (2018) prove convergence for general non-convex functions assuming gradients are bounded and the signs of stochastic gradients are the same along the trajectory. The analysis in (D'efossez et al., 2020) also relies on the bounded gradient assumption. Guo et al. (2021) assume the adaptive stepsize is upper and lower bounded by two constants, which is not necessarily satisfied unless assuming bounded gradients or considering the AdaBound variant (Luo et al., 2019) . (Zhang et al., 2022; Wang et al., 2022) consider very weak assumptions. However, they show either 1) "convergence" only to some neighborhood of stationary points with a constant radius, unless assuming the strong growth condition; or 2) convergence to stationary points but with a slower rate. Variants of Adam. After Reddi et al. ( 2018 ) pointed out the non-convergence issue with Adam, various variants of Adam that can be proved to converge were proposed (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ced2f706-3fd3-4d9b-b68a-8fd68b36b5a1Cited by top-tier papers59
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li et al.NeurIPS 2024 · 149 citations
- Convex and Non-convex Optimization Under Generalized SmoothnessHaochuan Li, Jian Qian, Yi Tian, Alexander Rakhlin et al.NeurIPS 2023 · 93 citations
- DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent MethodAhmed Khaled, Konstantin Mishchenko, Chi JinNeurIPS 2023 · 49 citations
- Muon Outperforms Adam in Tail-End Associative Memory LearningShuche Wang, Fengzhuo Zhang, Jiaxiang Li, Cunxiao Du et al.ICLR 2026 · 40 citations
- On Convergence of Adam for Stochastic Optimization under Relaxed AssumptionsYusu Hong, Junhong LinNeurIPS 2024 · 37 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex OptimizationZhize Li, Hongyan Bao, Xiangliang Zhang, Peter RichtárikICML 2021 · 164 citations
Related papers
- A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGDRuinan Jin, Xiao Li, Yaoliang Yu, Baoxiang WangICML 2025
- ADOPT: Modified Adam Can Converge with Any β2 with the Optimal RateShohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima et al.NeurIPS 2024 · 32 citations
- Closing the gap between the upper bound and lower bound of Adam's iteration complexityBohan Wang, Jingwen Fu, Huishuai Zhang, Nanning Zheng et al.NeurIPS 2023 · 2 citations
- Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and MomentumZeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato et al.ICML 2022 · 65 citations
- Towards Understanding Adam Convergence on Highly Degenerate PolynomialsZhiwei Bai, Jiajie Zhao, Zhangchen Zhou, Zhi-Qin John Xu et al.ICML 2026 · 2 citations
