Implicit Jacobian regularization weighted with impurity of probability output
Sungyoon Lee, Jinseong Park, Jaewook Lee
Abstract
The success of deep learning is greatly attributed to stochastic gradient descent (SGD), yet it remains unclear how SGD finds well-generalized models. We demonstrate that SGD has an implicit regularization effect on the logit-weight Jacobian norm of neural networks. This regularization effect is weighted with the impurity of the probability output, and thus it is active in a certain phase of training. Moreover, based on these findings, we propose a novel optimization method that explicitly regularizes the Jacobian norm, which leads to similar performance as other state-of-the-art sharpness-aware optimization methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Differentially Private Sharpness-Aware TrainingJinseong Park, Hoki Kim, Yujin Choi, Jaewook LeeICML 2023 · 15 citations
- Disentangling the Mechanisms Behind Implicit Regularization in SGDZachary Novack, Simran Kaur, Tanya Marwah, Saurabh Garg et al.ICLR 2023
- Rapid Selection and Ordering of In-Context Demonstrations via Prompt Embedding ClusteringKha Pham, Hung Le, Man Ngo, Truyen TranICLR 2025
- Quantifying and Optimizing Simplicity via Polynomial RepresentationsTianren Zhang, Xiangxin Li, Minghao Xiao, Guanyu Chen et al.ICML 2026
- Improving Adversarial Robustness of Attribution via Implicit RegularizationAmir Mehrpanah, Matteo Gamba, Hossein AzizpourICML 2026
Builds on10
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Bag of Tricks for Adversarial TrainingTianyu Pang, Xiao Yang, Yinpeng Dong, Hang Su et al.ICLR 2021 · 298 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 235 citations
Related papers
- Why is SAM Robust to Label Noise?Christina Baek, J. Zico Kolter, Aditi RaghunathanICLR 2024 · 24 citations
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 3 citations
- The Implicit and Explicit Regularization Effects of DropoutColin Wei, Sham M. Kakade, Tengyu MaICML 2020 · 129 citations
- The Implicit Regularization of Dynamical Stability in Stochastic Gradient DescentLei Wu, Weijie J. SuICML 2023 · 41 citations
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
