Neglected Hessian component explains mysteries in sharpness regularization
Yann N. Dauphin, Atish Agarwala, Hossein Mobahi
Abstract
Recent work has shown that methods like SAM which either explicitly or implicitly penalize second order information can improve generalization in deep learning. Seemingly similar methods like weight noise and gradient penalties often fail to provide such benefits. We show that these differences can be explained by the structure of the Hessian of the loss. First, we show that a common decomposition of the Hessian can be quantitatively interpreted as separating the feature exploitation from feature exploration. The feature exploration, which can be described by the Nonlinear Modeling Error matrix (NME), is commonly neglected in the literature since it vanishes at interpolation. Our work shows that the NME is in fact important as it can explain why gradient penalties are sensitive to the choice of activation function. Using this insight we design interventions to improve performance. We also provide evidence that challenges the long held equivalence of weight noise and gradient penalties. This equivalence relies on the assumption that the NME can be ignored, which we find does not hold for modern networks since they involve significant feature learning. We find that regularizing feature exploitation but not feature exploration yields performance similar to gradient penalties.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18ed7c34-ce8e-4821-befd-9571e5a30ef1Cited by top-tier papers9
- μP2: Effective Sharpness Aware Minimization Requires Layerwise Perturbation ScalingMoritz Haas, Jin Xu, Volkan Cevher, Leena Chennuru VankadaraNeurIPS 2024 · 12 citations
- Dual-Space Smoothness for Robust and Balanced LLM UnlearningHan Yan, Zheyuan Liu, Meng JiangICLR 2026 · 5 citations
- Leveraging Machine Unlearning for Cost-Efficient Preference AlignmentXiaoHua Feng, Yuyuan Li, HuWei Ji, Li Zhang et al.ICML 2026 · 4 citations
- Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and BeyondChongyu Fan, Jinghan Jia, Yihua Zhang, Anil Ramakrishna et al.ICML 2025
- LensLLM: Unveiling Fine-Tuning Dynamics for LLM SelectionXinyue Zeng, Haohui Wang, Junhong Lin, Jun Wu et al.ICML 2025
Builds on12
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 235 citations
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 190 citations
- Penalizing Gradient Norm for Efficiently Improving Generalization in Deep LearningYang Zhao, Hao Zhang, Xiuyuan HuICML 2022 · 165 citations
Related papers
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 3 citations
- Friendly Sharpness-Aware MinimizationTao Li, Pan Zhou, Zhengbao He, Xinwen Cheng et al.CVPR 2024
- Why Does Sharpness-Aware Minimization Generalize Better Than SGD?Zixiang Chen, Junkai Zhang, Yiwen Kou, Xiangning Chen et al.NeurIPS 2023 · 32 citations
- Layer-Wise Adaptive Gradient Norm Penalizing Method for Efficient and Accurate Deep LearningSunwoo LeeKDD 2024 · 2 citations
- Triple descent and the two kinds of overfitting: where & why do they appear?Stéphane d'Ascoli, Levent Sagun, Giulio BiroliNeurIPS 2020 · 94 citations
