Neglected Hessian component explains mysteries in sharpness regularization
Yann N. Dauphin, Atish Agarwala, Hossein Mobahi
摘要
Recent work has shown that methods like SAM which either explicitly or implicitly penalize second order information can improve generalization in deep learning. Seemingly similar methods like weight noise and gradient penalties often fail to provide such benefits. We show that these differences can be explained by the structure of the Hessian of the loss. First, we show that a common decomposition of the Hessian can be quantitatively interpreted as separating the feature exploitation from feature exploration. The feature exploration, which can be described by the Nonlinear Modeling Error matrix (NME), is commonly neglected in the literature since it vanishes at interpolation. Our work shows that the NME is in fact important as it can explain why gradient penalties are sensitive to the choice of activation function. Using this insight we design interventions to improve performance. We also provide evidence that challenges the long held equivalence of weight noise and gradient penalties. This equivalence relies on the assumption that the NME can be ignored, which we find does not hold for modern networks since they involve significant feature learning. We find that regularizing feature exploitation but not feature exploration yields performance similar to gradient penalties.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- μP2: Effective Sharpness Aware Minimization Requires Layerwise Perturbation ScalingMoritz Haas, Jin Xu, Volkan Cevher, Leena Chennuru VankadaraNeurIPS 2024 · 被引用 12 次
- Dual-Space Smoothness for Robust and Balanced LLM UnlearningHan Yan, Zheyuan Liu, Meng JiangICLR 2026 · 被引用 5 次
- Leveraging Machine Unlearning for Cost-Efficient Preference AlignmentXiaoHua Feng, Yuyuan Li, HuWei Ji, Li Zhang 等ICML 2026 · 被引用 4 次
- Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and BeyondChongyu Fan, Jinghan Jia, Yihua Zhang, Anil Ramakrishna 等ICML 2025
- LensLLM: Unveiling Fine-Tuning Dynamics for LLM SelectionXinyue Zeng, Haohui Wang, Junhong Lin, Jun Wu 等ICML 2025
它引用的顶会 Paper12
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 被引用 235 次
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 被引用 235 次
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 被引用 190 次
- Penalizing Gradient Norm for Efficiently Improving Generalization in Deep LearningYang Zhao, Hao Zhang, Xiuyuan HuICML 2022 · 被引用 165 次
相关 Paper
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 被引用 3 次
- Friendly Sharpness-Aware MinimizationTao Li, Pan Zhou, Zhengbao He, Xinwen Cheng 等CVPR 2024
- Why Does Sharpness-Aware Minimization Generalize Better Than SGD?Zixiang Chen, Junkai Zhang, Yiwen Kou, Xiangning Chen 等NeurIPS 2023 · 被引用 32 次
- Layer-Wise Adaptive Gradient Norm Penalizing Method for Efficient and Accurate Deep LearningSunwoo LeeKDD 2024 · 被引用 2 次
- Triple descent and the two kinds of overfitting: where & why do they appear?Stéphane d'Ascoli, Levent Sagun, Giulio BiroliNeurIPS 2020 · 被引用 94 次
