SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
Dmitry Kovalev
摘要
In this paper, we revisit stochastic gradient descent (SGD) with AdaGrad-type preconditioning. Our contributions are twofold. First, we develop a unified convergence analysis of SGD with adaptive preconditioning under anisotropic or matrix smoothness and noise assumptions. This allows us to recover state-of-the-art convergence results for several popular adaptive gradient methods, including AdaGrad-Norm, AdaGrad, and ASGO/One-sided Shampoo. In addition, we establish the fundamental connection between two recently proposed algorithms, Scion and DASGO, and provide the first theoretical guarantees for the latter. Second, we show that the convergence of methods like AdaGrad and DASGO can be provably accelerated beyond the best-known rates using Nesterov momentum. Consequently, we obtain the first theoretical justification that AdaGrad-type algorithms can simultaneously benefit from both diagonal preconditioning and momentum, which may provide an ultimate explanation for the practical efficiency of Adam.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex OptimizationEkaterina Borodich, Dmitry KovalevICLR 2026 · 被引用 10 次
- A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean DescentShuo Xie, Tianhao Wang, Beining Wu, Zhiyuan LiICLR 2026 · 被引用 7 次
它引用的顶会 Paper12
- ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally!Konstantin Mishchenko, Grigory Malinovsky, Sebastian U. Stich, Peter RichtárikICML 2022 · 被引用 200 次
- Prodigy: An Expeditiously Adaptive Parameter-Free LearnerKonstantin Mishchenko, Aaron DefazioICML 2024 · 被引用 131 次
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 被引用 117 次
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 被引用 98 次
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren 等NeurIPS 2025 · 被引用 58 次
相关 Paper
- Structured Preconditioners in Adaptive Optimization: A Unified AnalysisShuo Xie, Tianhao Wang, Sashank J. Reddi, Sanjiv Kumar 等ICML 2025
- Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided PreconditioningHuan Li, Yiming Dong, Zhouchen LinICML 2026
- Combining Axes Preconditioners through Kronecker Approximation for Deep LearningSai Surya Duvvuri, Devvrit, Rohan Anil, Cho-Jui Hsieh 等ICLR 2024 · 被引用 16 次
- Gradient correlation is a key ingredient to accelerate SGD with momentumJulien Hermant, Marien Renaud, Jean-François Aujol, Charles Dossal 等ICLR 2025
- A New Perspective on Shampoo's PreconditionerDepen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach 等ICLR 2025
