Normalization Layers Are All That Sharpness-Aware Minimization Needs
Maximilian Müller, Tiffany Vlaar, David Rolnick, Matthias Hein
Abstract
Sharpness-aware minimization (SAM) was proposed to reduce sharpness of minima and has been shown to enhance generalization performance in various settings. In this work we show that perturbing only the affine normalization parameters (typically comprising 0.1% of the total parameters) in the adversarial step of SAM can outperform perturbing all of the parameters. This finding generalizes to different SAM variants and both ResNet (Batch Normalization) and Vision Transformer (Layer Normalization) architectures. We consider alternative sparse perturbation approaches and find that these do not achieve similar performance enhancement at such extreme sparsity levels, showing that this behaviour is unique to the normalization layers. Although our findings reaffirm the effectiveness of SAM in improving generalization performance, they cast doubt on whether this is solely caused by reduced sharpness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a0ea459-a76b-41bc-83e8-eeb18fbcec25Cited by top-tier papers21
- Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware MinimizationZiqing Fan, Shengchao Hu, Jiangchao Yao, Gang Niu et al.ICML 2024 · 35 citations
- On the Duality Between Sharpness-Aware Minimization and Adversarial TrainingYihao Zhang, Hangzhou He, Jingyu Zhu, Huanran Chen et al.ICML 2024 · 29 citations
- Layer-wise linear mode connectivityLinara Adilova, Maksym Andriushchenko, Michael Kamp, Asja Fischer et al.ICLR 2024 · 22 citations
- Improving Generalization and Convergence by Enhancing Implicit RegularizationMingze Wang, Jinbo Wang, Haotian He, Zilin Wang et al.NeurIPS 2024 · 21 citations
- Momentum-SAM: Sharpness Aware Minimization without Computational OverheadMarlon Becker, Frederick Altrock, Benjamin RisseNeurIPS 2025 · 16 citations
Builds on27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
Related papers
- Layer-Wise Adaptive Gradient Norm Penalizing Method for Efficient and Accurate Deep LearningSunwoo LeeKDD 2024 · 2 citations
- Sharpness-Aware Minimization Leads to Low-Rank FeaturesMaksym Andriushchenko, Dara Bahri, Hossein Mobahi, Nicolas FlammarionNeurIPS 2023 · 48 citations
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 3 citations
- Friendly Sharpness-Aware MinimizationTao Li, Pan Zhou, Zhengbao He, Xinwen Cheng et al.CVPR 2024
- Beyond Sharpness: The Role of Nonuniformity in GeneralizationYingcong Zhou, Pingfan Wu, Li Wang, Zhiguo Fu et al.AAAI 2026
