A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
Ruinan Jin, Xiao Li, Yaoliang Yu, Baoxiang Wang
摘要
Adaptive moment estimation (Adam) is a cornerstone optimization algorithm in deep learning, widely recognized for its flexibility with adaptive learning rates and efficiency in handling largescale data. However, despite its practical success, the theoretical understanding of Adam's convergence has been constrained by stringent assumptions, such as almost surely bounded stochastic gradients or uniformly bounded gradients, which are more restrictive than those typically required for analyzing stochastic gradient descent (SGD). In this paper, we introduce a novel and comprehensive framework for analyzing the convergence properties of Adam. This framework offers a versatile approach to establishing Adam's convergence. Specifically, we prove that Adam achieves asymptotic (last iterate sense) convergence in both the almost sure sense and the L 1 sense under the relaxed assumptions typically used for SGD, namely L-smoothness and the ABC inequality. Meanwhile, under the same assumptions, we show that Adam attains non-asymptotic sample complexity bounds similar to those of SGD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- On the O(√d/K1/4) Convergence Rate of AdamW Measured by ℓ1 NormHuan Li, Yiming Dong, Zhouchen LinNeurIPS 2025 · 被引用 10 次
- Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch RegimeBeomhan Baek, Minhak Song, Chulhee YunICLR 2026 · 被引用 2 次
它引用的顶会 Paper4
- An Improved Analysis of Stochastic Gradient Descent with MomentumYanli Liu, Yuan Gao, Wotao YinNeurIPS 2020 · 被引用 328 次
- Convergence of Adam Under Relaxed AssumptionsHaochuan Li, Alexander Rakhlin, Ali JadbabaieNeurIPS 2023 · 被引用 132 次
- SUPER-ADAM: Faster and Universal Framework of Adaptive GradientsFeihu Huang, Junyi Li, Heng HuangNeurIPS 2021 · 被引用 55 次
- Closing the gap between the upper bound and lower bound of Adam's iteration complexityBohan Wang, Jingwen Fu, Huishuai Zhang, Nanning Zheng 等NeurIPS 2023 · 被引用 2 次
相关 Paper
- Provable Adaptivity of Adam under Non-uniform SmoothnessBohan Wang, Yushun Zhang, Huishuai Zhang, Qi Meng 等KDD 2024 · 被引用 4 次
- On Convergence of Adam for Stochastic Optimization under Relaxed AssumptionsYusu Hong, Junhong LinNeurIPS 2024 · 被引用 37 次
- ADOPT: Modified Adam Can Converge with Any β2 with the Optimal RateShohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima 等NeurIPS 2024 · 被引用 32 次
- Robustness to Unbounded Smoothness of Generalized SignSGDMichael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang 等NeurIPS 2022 · 被引用 111 次
- Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and MomentumZeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato 等ICML 2022 · 被引用 65 次
