A new regret analysis for Adam-type algorithms
Ahmet Alacaoglu, Yura Malitsky, Panayotis Mertikopoulos, Volkan Cevher
2020年份
50被引次数
20顶会引用
摘要
In this paper, we focus on a theory-practice gap for Adam and its variants (AMSgrad, AdamNC, etc.). In practice, these algorithms are used with a constant first-order moment parameter (typically between and ). In theory, regret guarantees for online convex optimization require a rapidly decaying schedule. We show that this is an artifact of the standard analysis and propose a novel framework that allows us to derive optimal, data-dependent regret bounds with a constant , without further assumptions. We also demonstrate the flexibility of our analysis on a wide range of different algorithms and settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- 1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence SpeedHanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari 等ICML 2021 · 被引用 106 次
- High Probability Bounds for a Class of Nonconvex Algorithms with AdaGrad StepsizeAli Kavis, Kfir Yehuda Levy, Volkan CevherICLR 2022 · 被引用 51 次
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 被引用 43 次
- Adam with model exponential moving average is effective for nonconvex optimizationKwangjun Ahn, Ashok CutkoskyNeurIPS 2024 · 被引用 36 次
- SGD with AdaGrad Stepsizes: Full Adaptivity with High Probability to Unknown Parameters, Unbounded Gradients and Affine VarianceAmit Attia, Tomer KorenICML 2023 · 被引用 34 次
它引用的顶会 Paper1
相关 Paper
- Between Stochastic and Adversarial Online Convex Optimization: Improved Regret Bounds via SmoothnessSarah Sachs, Hédi Hadiji, Tim van Erven, Cristóbal GuzmánNeurIPS 2022 · 被引用 30 次
- Better Full-Matrix Regret via Parameter-Free Online LearningAshok CutkoskyNeurIPS 2020 · 被引用 7 次
- Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam OptimizerYan-Feng Xie, Yu-Jie Zhang, Peng Zhao, Zhi-Hua ZhouICML 2026 · 被引用 2 次
- Closing the gap between the upper bound and lower bound of Adam's iteration complexityBohan Wang, Jingwen Fu, Huishuai Zhang, Nanning Zheng 等NeurIPS 2023 · 被引用 2 次
- Towards Understanding Adam Convergence on Highly Degenerate PolynomialsZhiwei Bai, Jiajie Zhao, Zhangchen Zhou, Zhi-Qin John Xu 等ICML 2026 · 被引用 2 次
