Implicit regularization in Heavy-ball momentum accelerated stochastic gradient descent
Avrajit Ghosh, He Lyu, Xitong Zhang, Rongrong Wang
Abstract
It is well known that the finite step-size () in Gradient Descent (GD) implicitly regularizes solutions to flatter minima. A natural question to ask is "Does the momentum parameter play a role in implicit regularization in Heavy-ball (H.B) momentum accelerated gradient descent (GD+M)?". To answer this question, first, we show that the discrete H.B momentum update (GD+M) follows a continuous trajectory induced by a modified loss, which consists of an original loss and an implicit regularizer. Then, we show that this implicit regularizer for (GD+M) is stronger than that of (GD) by factor of , thus explaining why (GD+M) shows better generalization performance and higher test accuracy than (GD). Furthermore, we extend our analysis to the stochastic version of gradient descent with momentum (SGD+M) and characterize the continuous trajectory of the update of (SGD+M) in a pointwise sense. We explore the implicit regularization in (SGD+M) and (GD+M) through a series of experiments validating our theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efce7de0-97d3-4c71-93eb-e0a95145b330Cited by top-tier papers13
- On the Implicit Bias of AdamMatias D. Cattaneo, Jason M. Klusowski, Boris ShigidaICML 2024 · 26 citations
- The Marginal Value of Momentum for Small Learning Rate SGDRunzhe Wang, Sadhika Malladi, Tianhao Wang, Kaifeng Lyu et al.ICLR 2024 · 14 citations
- The Impact of Geometric Complexity on Neural Collapse in Transfer LearningMichael Munn, Benoit Dherin, Javier GonzalvoNeurIPS 2024 · 6 citations
- Where Do Large Learning Rates Lead Us?Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny, Ekaterina Lobacheva et al.NeurIPS 2024 · 6 citations
- How Memory in Optimization Algorithms Implicitly Modifies the LossMatias D. Cattaneo, Boris ShigidaNeurIPS 2025 · 6 citations
Builds on9
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 178 citations
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan et al.ICML 2020 · 125 citations
- Explicit Regularisation in Gaussian Noise InjectionsAlexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J. Roberts et al.NeurIPS 2020 · 90 citations
Related papers
- Effects of Momentum in Implicit Bias of Gradient Flow for Diagonal Linear NetworksBochen Lyu, He Wang, Zheng Wang, Zhanxing ZhuAAAI 2025 · 1 citation
- Heavy-Ball Momentum Method in Continuous Time and Discretization Error AnalysisBochen Lyu, Xiaojing Zhang, Fangyi Zheng, He Wang et al.NeurIPS 2025 · 1 citation
- Does Momentum Change the Implicit Regularization on Separable Data?Bohan Wang, Qi Meng, Huishuai Zhang, Ruoyu Sun et al.NeurIPS 2022 · 29 citations
- Continuous-Time Analysis of Heavy Ball Momentum in Min-Max GamesYi Feng, Kaito Fujii, Stratis Skoulakis, Xiao Wang et al.ICML 2025
- Towards understanding how momentum improves generalization in deep learningSamy Jelassi, Yuanzhi LiICML 2022 · 53 citations
