Lune

KDD2026顶会

Adaptive Momentum by Momentum for Deep Neural Network Training

Tao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li, Bao Wang

2026年份
1被引次数

摘要

Heavy-ball momentum accelerates gradient-based optimization methods in machine learning. A standard way to use momentum is to choose a fixed hyperparameter, but this choice often requires extensive tuning. Relying on a one-size-fits-all momentum parameter can prevent an optimizer from achieving its best performance. Motivated by the optimal heavy-ball momentum for quadratic objectives, this paper proposes a new adaptive momentum mechanism that reduces the burden of momentum tuning. The proposed mechanism improves the stability of SGD and Adam under large learning rates, leading to faster convergence and better generalization than standard SGD and Adam in our experiments. We demonstrate the effectiveness of the method on a wide range of machine learning benchmarks, including image classification, language modeling, machine translation, tabular learning, and time-series classification. We also provide convergence guarantees for SGD and Adam equipped with the proposed adaptive momentum.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 3ffeb85e-0340-4df5-90d2-e425644a0f24

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖