Adaptive Momentum by Momentum for Deep Neural Network Training
Tao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li, Bao Wang
摘要
Heavy-ball momentum accelerates gradient-based optimization methods in machine learning. A standard way to use momentum is to choose a fixed hyperparameter, but this choice often requires extensive tuning. Relying on a one-size-fits-all momentum parameter can prevent an optimizer from achieving its best performance. Motivated by the optimal heavy-ball momentum for quadratic objectives, this paper proposes a new adaptive momentum mechanism that reduces the burden of momentum tuning. The proposed mechanism improves the stability of SGD and Adam under large learning rates, leading to faster convergence and better generalization than standard SGD and Adam in our experiments. We demonstrate the effectiveness of the method on a wide range of machine learning benchmarks, including image classification, language modeling, machine translation, tabular learning, and time-series classification. We also provide convergence guarantees for SGD and Adam equipped with the proposed adaptive momentum.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Demystify Hyperparameters for Stochastic Optimization with Transferable RepresentationsJianhui Sun, Mengdi Huai, Kishlay Jha, Aidong ZhangKDD 2022 · 被引用 5 次
- Provable Acceleration of Heavy Ball beyond Quadratics for a Class of Polyak-Lojasiewicz Functions when the Non-Convexity is Averaged-OutJun-Kun Wang, Chi-Heng Lin, Andre Wibisono, Bin HuICML 2022 · 被引用 27 次
- A Stagewise Hyperparameter Scheduler to Improve GeneralizationJianhui Sun, Ying Yang, Guangxu Xun, Aidong ZhangKDD 2021 · 被引用 8 次
- Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient NoiseRui Pan, Yuxing Liu, Xiaoyu Wang, Tong ZhangICLR 2024 · 被引用 10 次
- Generalized Polyak Step Size for First Order Optimization with MomentumXiaoyu Wang, Mikael Johansson, Tong ZhangICML 2023 · 被引用 32 次
