Domain-Independent Dominance of Adaptive Methods
Pedro Savarese, David McAllester, Sudarshan Babu, Michael Maire
摘要
From a simplified analysis of adaptive methods, we derive AvaGrad, a new optimizer which outperforms SGD on vision tasks when its adaptability is properly tuned. We observe that the power of our method is partially explained by a decoupling of learning rate and adaptability, greatly simplifying hyperparameter search. In light of this observation, we demonstrate that, against conventional wisdom, Adam can also outperform SGD on vision tasks, as long as the coupling between its learning rate and adaptability is taken into account. In practice, AvaGrad matches the best results, as measured by generalization accuracy, delivered by any existing optimizer (SGD or adaptive) across image classification (CIFAR, ImageNet) and character-level language modelling (Penn Treebank) tasks. When training GANs, AvaGrad improves upon existing optimizers. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 被引用 195 次
- Permutation Search of Tensor Network Structures via Local SamplingChao Li, Junhua Zeng, Zerui Tao, Qibin ZhaoICML 2022 · 被引用 31 次
- On the O(√d/K1/4) Convergence Rate of AdamW Measured by ℓ1 NormHuan Li, Yiming Dong, Zhouchen LinNeurIPS 2025 · 被引用 10 次
- Momentum Centering and Asynchronous Update for Adaptive Gradient MethodsJuntang Zhuang, Yifan Ding, Tommy Tang, Nicha C. Dvornek 等NeurIPS 2021 · 被引用 9 次
- Accelerated Training via Incrementally Growing Neural Networks using Variance Transfer and Learning Rate AdaptationXin Yuan, Pedro Savarese, Michael MaireNeurIPS 2023 · 被引用 9 次
它引用的顶会 Paper5
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda 等NeurIPS 2020 · 被引用 697 次
- ContraGAN: Contrastive Learning for Conditional Image GenerationMinguk Kang, Jaesik ParkNeurIPS 2020 · 被引用 216 次
- Gradientless Descent: High-Dimensional Zeroth-Order OptimizationDaniel Golovin, John Karro, Greg Kochanski, Chansoo Lee 等ICLR 2020 · 被引用 85 次
- On the Adequacy of Untuned Warmup for Adaptive OptimizationJerry Ma, Denis YaratsAAAI 2021 · 被引用 81 次
相关 Paper
- MADA: Meta-Adaptive Optimizers Through Hyper-Gradient DescentKaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong 等ICML 2024 · 被引用 7 次
- Gradient descent with generalized Newton's methodZhiqi Bu, Shiyun XuICLR 2025
- HVAdam: A Full-Dimension Adaptive OptimizerYiheng Zhang, Shaowu Wu, Yuanzhuo Xu, Jiajun Wu 等AAAI 2025 · 被引用 1 次
- ADOPT: Modified Adam Can Converge with Any β2 with the Optimal RateShohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima 等NeurIPS 2024 · 被引用 32 次
- AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning MatrixYun Yue, Zhiling Ye, Jiadi Jiang, Yongchao Liu 等NeurIPS 2023 · 被引用 6 次
