Avoiding spurious sharpness minimization broadens applicability of SAM
Sidak Pal Singh, Hossein Mobahi, Atish Agarwala, Yann N. Dauphin
Abstract
Curvature regularization techniques like Sharpness Aware Minimization (SA M) have shown great promise in improving generalization on vision tasks. However, we find that SA M performs poorly in domains like natural language processing (NLP), often degrading performance -even with twice the compute budget. We investigate the discrepancy across domains and find that in the NLP setting, SAM is dominated by regularization of the logit statistics --instead of improving the geometry of the function itself. We use this observation to develop an alternative algorithm we call Fu nc t i onal -SA M, which regularizes curvature only through modification of the statistics of the overall function implemented by the neural network, and avoids spurious minimization through logit manipulation. Furthermore, we argue that preconditioning the SA M perturbation also prevents spurious minimization, and when combined with Fu nc t i onal -SA M, it gives further improvements. Our proposed algorithms show improved performance over A dam W and SA M baselines when trained for an equal number of steps, in both fixed-length and Chinchilla-style training settings, at various model scales (including billionparameter scale). On the whole, our work highlights the importance of more precise characterizations of sharpness in broadening the applicability of curvature regularization to large language models (LLMs).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f0c3cb5-02e3-4f09-86e3-0acfbca98aa6Cited by top-tier papers6
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient NoiseHaocheng Luo, Mehrtash Harandi, Dinh Phung, Trung LeNeurIPS 2025 · 2 citations
- Flat Minima and Generalization: Insights from Stochastic Convex OptimizationMatan Schliserman, Shira Vansover-Hager, Tomer KorenICML 2026 · 2 citations
- Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference OptimizationHaocheng Luo, Zehang Deng, Thanh-Toan Do, Mehrtash Harandi et al.ICLR 2026 · 1 citation
- Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded SchedulerDimitris Oikonomou, Nicolas LoizouICML 2026
- Sharpness-Aware Minimization: General Analysis and Improved RatesDimitris Oikonomou, Nicolas LoizouICLR 2025
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 385 citations
Related papers
- TRAM: Bridging Trust Regions and Sharpness Aware MinimizationTom Sherborne, Naomi Saphra, Pradeep Dasigi, Hao PengICLR 2024 · 6 citations
- GA-SAM: Gradient-Strength based Adaptive Sharpness-Aware Minimization for Improved GeneralizationZhiyuan Zhang, Ruixuan Luo, Qi Su, Xu SunEMNLP 2022 · 9 citations
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 3 citations
- Sharpness-Aware Minimization Improves Language Model GeneralizationDara Bahri, Hossein Mobahi, Yi TayACL 2022
- Fix the Loss, Not the Radius: Rethinking the Adversarial Perturbation of Sharpness-Aware MinimizationJinping Wang, Qinhan Liu, Zhiwu Xie, Zhiqiang GaoICML 2026
