Adaptive Momentum by Momentum for Deep Neural Network Training
Tao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li, Bao Wang
Abstract
Heavy-ball momentum accelerates gradient-based optimization methods in machine learning. A standard way to use momentum is to choose a fixed hyperparameter, but this choice often requires extensive tuning. Relying on a one-size-fits-all momentum parameter can prevent an optimizer from achieving its best performance. Motivated by the optimal heavy-ball momentum for quadratic objectives, this paper proposes a new adaptive momentum mechanism that reduces the burden of momentum tuning. The proposed mechanism improves the stability of SGD and Adam under large learning rates, leading to faster convergence and better generalization than standard SGD and Adam in our experiments. We demonstrate the effectiveness of the method on a wide range of machine learning benchmarks, including image classification, language modeling, machine translation, tabular learning, and time-series classification. We also provide convergence guarantees for SGD and Adam equipped with the proposed adaptive momentum.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3ffeb85e-0340-4df5-90d2-e425644a0f24Related papers
- Demystify Hyperparameters for Stochastic Optimization with Transferable RepresentationsJianhui Sun, Mengdi Huai, Kishlay Jha, Aidong ZhangKDD 2022 · 5 citations
- Provable Acceleration of Heavy Ball beyond Quadratics for a Class of Polyak-Lojasiewicz Functions when the Non-Convexity is Averaged-OutJun-Kun Wang, Chi-Heng Lin, Andre Wibisono, Bin HuICML 2022 · 27 citations
- A Stagewise Hyperparameter Scheduler to Improve GeneralizationJianhui Sun, Ying Yang, Guangxu Xun, Aidong ZhangKDD 2021 · 8 citations
- Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient NoiseRui Pan, Yuxing Liu, Xiaoyu Wang, Tong ZhangICLR 2024 · 10 citations
- Generalized Polyak Step Size for First Order Optimization with MomentumXiaoyu Wang, Mikael Johansson, Tong ZhangICML 2023 · 32 citations
