Implicit regularization in Heavy-ball momentum accelerated stochastic gradient descent
Avrajit Ghosh, He Lyu, Xitong Zhang, Rongrong Wang
摘要
It is well known that the finite step-size () in Gradient Descent (GD) implicitly regularizes solutions to flatter minima. A natural question to ask is "Does the momentum parameter play a role in implicit regularization in Heavy-ball (H.B) momentum accelerated gradient descent (GD+M)?". To answer this question, first, we show that the discrete H.B momentum update (GD+M) follows a continuous trajectory induced by a modified loss, which consists of an original loss and an implicit regularizer. Then, we show that this implicit regularizer for (GD+M) is stronger than that of (GD) by factor of , thus explaining why (GD+M) shows better generalization performance and higher test accuracy than (GD). Furthermore, we extend our analysis to the stochastic version of gradient descent with momentum (SGD+M) and characterize the continuous trajectory of the update of (SGD+M) in a pointwise sense. We explore the implicit regularization in (SGD+M) and (GD+M) through a series of experiments validating our theory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- On the Implicit Bias of AdamMatias D. Cattaneo, Jason M. Klusowski, Boris ShigidaICML 2024 · 被引用 26 次
- The Marginal Value of Momentum for Small Learning Rate SGDRunzhe Wang, Sadhika Malladi, Tianhao Wang, Kaifeng Lyu 等ICLR 2024 · 被引用 14 次
- The Impact of Geometric Complexity on Neural Collapse in Transfer LearningMichael Munn, Benoit Dherin, Javier GonzalvoNeurIPS 2024 · 被引用 6 次
- Where Do Large Learning Rates Lead Us?Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny, Ekaterina Lobacheva 等NeurIPS 2024 · 被引用 6 次
- How Memory in Optimization Algorithms Implicitly Modifies the LossMatias D. Cattaneo, Boris ShigidaNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper9
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 被引用 235 次
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 被引用 178 次
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan 等ICML 2020 · 被引用 125 次
- Explicit Regularisation in Gaussian Noise InjectionsAlexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J. Roberts 等NeurIPS 2020 · 被引用 90 次
相关 Paper
- Effects of Momentum in Implicit Bias of Gradient Flow for Diagonal Linear NetworksBochen Lyu, He Wang, Zheng Wang, Zhanxing ZhuAAAI 2025 · 被引用 1 次
- Heavy-Ball Momentum Method in Continuous Time and Discretization Error AnalysisBochen Lyu, Xiaojing Zhang, Fangyi Zheng, He Wang 等NeurIPS 2025 · 被引用 1 次
- Does Momentum Change the Implicit Regularization on Separable Data?Bohan Wang, Qi Meng, Huishuai Zhang, Ruoyu Sun 等NeurIPS 2022 · 被引用 29 次
- Continuous-Time Analysis of Heavy Ball Momentum in Min-Max GamesYi Feng, Kaito Fujii, Stratis Skoulakis, Xiao Wang 等ICML 2025
- Towards understanding how momentum improves generalization in deep learningSamy Jelassi, Yuanzhi LiICML 2022 · 被引用 53 次
