Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias
Ryo Karakida, Tomoumi Takase, Tomohiro Hayase, Kazuki Osawa
Abstract
Gradient regularization (GR) is a method that penalizes the gradient norm of the training loss during training. While some studies have reported that GR can improve generalization performance, little attention has been paid to it from the algorithmic perspective, that is, the algorithms of GR that efficiently improve the performance. In this study, we first reveal that a specific finite-difference computation, composed of both gradient ascent and descent steps, reduces the computational cost of GR. Next, we show that the finite-difference computation also works better in the sense of generalization performance. We theoretically analyze a solvable model, a diagonal linear network, and clarify that GR has a desirable implicit bias to so-called rich regime and finite-difference computation strengthens this bias. Furthermore, finite-difference GR is closely related to some other algorithms based on iterative ascent and descent steps for exploring flat minima. In particular, we reveal that the flooding method can perform finite-difference GR in an implicit way. Thus, this work broadens our understanding of GR for both practice and theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c13c861-4b71-4b83-b28e-4c5c13e9a2ddCited by top-tier papers10
- Adam Reduces a Unique Form of Sharpness: Theoretical Insights Near the Minimizer ManifoldXinghan Li, Haodong Wen, Kaifeng LyuNeurIPS 2025 · 6 citations
- Soft ascent-descent as a stable and flexible alternative to floodingMatthew J. Holland, Kosuke NakataniNeurIPS 2024 · 4 citations
- When Will Gradient Regularization Be Harmful?Yang Zhao, Hao Zhang, Xiuyuan HuICML 2024 · 3 citations
- Criterion Collapse and Loss Distribution ControlMatthew J. HollandICML 2024 · 2 citations
- An Infinite-Width Analysis on the Jacobian-Regularised Training of a Neural NetworkTaeyoung Kim, Hongseok YangICML 2024 · 2 citations
Builds on13
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 235 citations
- Surrogate Gap Minimization Improves Sharpness-Aware TrainingJuntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui et al.ICLR 2022 · 213 citations
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 190 citations
Related papers
- Do We Need Zero Training Loss After Achieving Zero Training Error?Takashi Ishida, Ikko Yamane, Tomoya Sakai, Gang Niu et al.ICML 2020 · 155 citations
- Grokking Beyond the Euclidean Norm of Model ParametersPascal Tikeng Notsawo Jr., Guillaume Dumas, Guillaume RabusseauICML 2025
- On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning RegimeShuai Jiang, Eric C. Cyr, Ben S. Southworth, Alexey VoroninICLR 2026 · 1 citation
- Conflicting Biases at the Edge of Stability: Norm versus Sharpness RegularizationMaria Matveev, Vit Fojtik, Hung-Hsu Chou, Gitta Kutyniok et al.ICML 2026
- Implicit Bias of (Stochastic) Gradient Descent for Rank-1 Linear Neural NetworkBochen Lyu, Zhanxing ZhuNeurIPS 2023 · 5 citations
