Adam with model exponential moving average is effective for nonconvex optimization
Kwangjun Ahn, Ashok Cutkosky
Abstract
In this work, we offer a theoretical analysis of two modern optimization techniques for training large and complex models: (i) adaptive optimization algorithms, such as Adam, and (ii) the model exponential moving average (EMA). Specifically, we demonstrate that a clipped version of Adam with model EMA achieves the optimal convergence rates in various nonconvex optimization settings, both smooth and nonsmooth. Moreover, when the scale varies significantly across different coordinates, we demonstrate that the coordinate-wise adaptivity of Adam is provably advantageous. Notably, unlike previous analyses of Adam, our analysis crucially relies on its core elements -- momentum and discounting factors -- as well as model EMA, motivating their wide applications in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2acb2ad-8454-473e-a455-af63e4263c9fCited by top-tier papers10
- Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in DisguiseKwangjun Ahn, Zhiyu Zhang, Yunbum Kook, Yan DaiICML 2024 · 25 citations
- On the O(√d/K1/4) Convergence Rate of AdamW Measured by ℓ1 NormHuan Li, Yiming Dong, Zhouchen LinNeurIPS 2025 · 10 citations
- Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model TrainingMinhak Song, Beomhan Baek, Kwangjun Ahn, Chulhee YunNeurIPS 2025 · 9 citations
- Improving Online-to-Nonconvex Conversion for Smooth Optimization via Double OptimismFrancisco Patitucci, Ruichen Jiang, Aryan MokhtariICLR 2026 · 3 citations
- Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam OptimizerYan-Feng Xie, Yu-Jie Zhang, Peng Zhao, Zhi-Hua ZhouICML 2026 · 2 citations
Builds on26
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- The AdEMAMix Optimizer: Better, Faster, OlderMatteo Pagliardini, Pierre Ablin, David GrangierICLR 2025
- Momentum Improves Normalized SGDAshok Cutkosky, Harsh MehtaICML 2020 · 177 citations
- How to Scale Your EMADan Busbridge, Jason Ramapuram, Pierre Ablin, Tatiana Likhomanenko et al.NeurIPS 2023 · 33 citations
- How to set AdamW's weight decay as you scale model and dataset sizeXi Wang, Laurence AitchisonICML 2025
- A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGDRuinan Jin, Xiao Li, Yaoliang Yu, Baoxiang WangICML 2025
