Avoiding spurious sharpness minimization broadens applicability of SAM
Sidak Pal Singh, Hossein Mobahi, Atish Agarwala, Yann N. Dauphin
摘要
Curvature regularization techniques like Sharpness Aware Minimization (SA M) have shown great promise in improving generalization on vision tasks. However, we find that SA M performs poorly in domains like natural language processing (NLP), often degrading performance -even with twice the compute budget. We investigate the discrepancy across domains and find that in the NLP setting, SAM is dominated by regularization of the logit statistics --instead of improving the geometry of the function itself. We use this observation to develop an alternative algorithm we call Fu nc t i onal -SA M, which regularizes curvature only through modification of the statistics of the overall function implemented by the neural network, and avoids spurious minimization through logit manipulation. Furthermore, we argue that preconditioning the SA M perturbation also prevents spurious minimization, and when combined with Fu nc t i onal -SA M, it gives further improvements. Our proposed algorithms show improved performance over A dam W and SA M baselines when trained for an equal number of steps, in both fixed-length and Chinchilla-style training settings, at various model scales (including billionparameter scale). On the whole, our work highlights the importance of more precise characterizations of sharpness in broadening the applicability of curvature regularization to large language models (LLMs).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient NoiseHaocheng Luo, Mehrtash Harandi, Dinh Phung, Trung LeNeurIPS 2025 · 被引用 2 次
- Flat Minima and Generalization: Insights from Stochastic Convex OptimizationMatan Schliserman, Shira Vansover-Hager, Tomer KorenICML 2026 · 被引用 2 次
- Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference OptimizationHaocheng Luo, Zehang Deng, Thanh-Toan Do, Mehrtash Harandi 等ICLR 2026 · 被引用 1 次
- Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded SchedulerDimitris Oikonomou, Nicolas LoizouICML 2026
- Sharpness-Aware Minimization: General Analysis and Improved RatesDimitris Oikonomou, Nicolas LoizouICLR 2025
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 被引用 385 次
相关 Paper
- TRAM: Bridging Trust Regions and Sharpness Aware MinimizationTom Sherborne, Naomi Saphra, Pradeep Dasigi, Hao PengICLR 2024 · 被引用 6 次
- GA-SAM: Gradient-Strength based Adaptive Sharpness-Aware Minimization for Improved GeneralizationZhiyuan Zhang, Ruixuan Luo, Qi Su, Xu SunEMNLP 2022 · 被引用 9 次
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 被引用 3 次
- Sharpness-Aware Minimization Improves Language Model GeneralizationDara Bahri, Hossein Mobahi, Yi TayACL 2022
- Fix the Loss, Not the Radius: Rethinking the Adversarial Perturbation of Sharpness-Aware MinimizationJinping Wang, Qinhan Liu, Zhiwu Xie, Zhiqiang GaoICML 2026
