Lune

KDD2026Top-tier venue

Adaptive Momentum by Momentum for Deep Neural Network Training

Tao Sun, Huaming Ling, Zuoqiang Shi, Dongsheng Li, Bao Wang

2026Year
1Citations

Abstract

Heavy-ball momentum accelerates gradient-based optimization methods in machine learning. A standard way to use momentum is to choose a fixed hyperparameter, but this choice often requires extensive tuning. Relying on a one-size-fits-all momentum parameter can prevent an optimizer from achieving its best performance. Motivated by the optimal heavy-ball momentum for quadratic objectives, this paper proposes a new adaptive momentum mechanism that reduces the burden of momentum tuning. The proposed mechanism improves the stability of SGD and Adam under large learning rates, leading to faster convergence and better generalization than standard SGD and Adam in our experiments. We demonstrate the effectiveness of the method on a wide range of machine learning benchmarks, including image classification, language modeling, machine translation, tabular learning, and time-series classification. We also provide convergence guarantees for SGD and Adam equipped with the proposed adaptive momentum.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 3ffeb85e-0340-4df5-90d2-e425644a0f24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines