Towards Better Generalization of Adaptive Gradient Methods
Yingxue Zhou, Belhal Karimi, Jinxing Yu, Zhiqiang Xu, Ping Li
Abstract
Adaptive gradient methods such as AdaGrad, RMSprop and Adam have been optimizers of choice for deep learning due to their fast training speed. However, it was recently observed that their generalization performance is often worse than that of SGD for over-parameterized neural networks. While new algorithms (such as Ad-aBound) have been proposed to improve the situation, the provided analyses are only committed to optimization bounds for the training objective, leaving critical generalization capacity unexplored. To close this gap, we propose Stable Adaptive Gradient Descent (SAGD) for non-convex optimization which leverages differential privacy to boost the generalization performance of adaptive gradient methods. Theoretical analyses show that SAGD has high-probability convergence to a population stationary point. We further conduct experiments on various popular deep learning tasks and models. Experimental results illustrate that SAGD is empirically competitive and often better than baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Analysis of Error Feedback in Federated Non-Convex Optimization with Biased Compression: Fast Convergence and Partial ParticipationXiaoyun Li, Ping LiICML 2023 · 42 citations
- Generalization Guarantee of SGD for Pairwise LearningYunwen Lei, Mingrui Liu, Yiming YingNeurIPS 2021 · 37 citations
- On Distributed Adaptive Optimization with Gradient CompressionXiaoyun Li, Belhal Karimi, Ping LiICLR 2022 · 34 citations
- PoF: Post-Training of Feature Extractor for Improving GeneralizationIkuro Sato, Ryota Yamada, Masayuki Tanaka, Nakamasa Inoue et al.ICML 2022 · 5 citations
- Some Optimizers are More Equal: Understanding the Role of Optimizers in Group FairnessMojtaba Kolahdouzi, Hatice Gunes, Ali EtemadNeurIPS 2025
Builds on1
Related papers
- Understanding the Generalization of Adam in Learning Neural Networks with Proper RegularizationDifan Zou, Yuan Cao, Yuanzhi Li, Quanquan GuICLR 2023 · 6 citations
- Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and MomentumZeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato et al.ICML 2022 · 65 citations
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 25 citations
- Non-asymptotic Analysis of Biased Adaptive Stochastic ApproximationSobihan Surendran, Adeline Fermanian, Antoine Godichon-Baggioni, Sylvain Le CorffNeurIPS 2024 · 7 citations
- Improved Rates of Differentially Private Nonconvex-Strongly-Concave Minimax OptimizationRuijia Zhang, Mingxi Lei, Meng Ding, Zihang Xiang et al.AAAI 2025 · 7 citations
