Enhancing Sharpness-Aware Optimization Through Variance Suppression
Bingcong Li, Georgios B. Giannakis
Abstract
Sharpness-aware minimization (SAM) has well documented merits in enhancing generalization of deep neural networks, even without sizable data augmentation. Embracing the geometry of the loss function, where neighborhoods of 'flat minima' heighten generalization ability, SAM seeks 'flat valleys' by minimizing the maximum loss caused by an adversary perturbing parameters within the neighborhood. Although critical to account for sharpness of the loss function, such an 'over-friendly adversary' can curtail the outmost level of generalization. The novel approach of this contribution fosters stabilization of adversaries through variance suppression (VaSSO) to avoid such friendliness. VaSSO's provable stability safeguards its numerical improvement over SAM in model-agnostic tasks, including image classification and machine translation. In addition, experiments confirm that VaSSO endows SAM with robustness against high levels of label noise. Code is available at https://github.com/ BingcongLi/VaSSO . Introduction Despite deep neural networks (DNNs) have advanced the concept of "learning from data," and markedly improved performance across several applications in vision and language (Devlin et al., 2018; Tom et al., 2020) , their overparametrized nature renders the tendency to overfit on training data (Zhang et al., 2021a). This has led to concerns in generalization, which is a practically underscored perspective yet typically suffers from a gap relative to the training performance. Improving generalizability is challenging. Common approaches include (model) regularization and data augmentation (Srivastava et al., 2014) . While it is the default choice to integrate regularization such as weight decay and dropout into training, these methods are often insufficient for DNNs especially when coping with complicated network architectures (Chen et al., 2022) . Another line of effort resorts to suitable optimization schemes attempting to find a generalizable local minimum. For example, SGD is more preferable than Adam on certain overparameterized problems since it converges to maximum margin solutions (Wilson et al., 2017) . Decoupling weight decay from Adam also empirically facilitates generalizability (Loshchilov and Hutter, 2017) . Unfortunately, the underlying mechanism remains unveiled, and whether the generalization merits carry over to other intricate learning tasks calls for additional theoretical elaboration. Our main focus, sharpness aware minimization (SAM), is a highly compelling optimization approach that facilitates state-of-the-art generalizability by exploiting sharpness of loss landscape (Foret et al., 2021; Chen et al., 2022) . A high-level interpretation of sharpness is how violently the loss fluctuates within a neighborhood. It has been shown through large-scale empirical studies that sharpness-based measures highly
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78e47a7c-6059-4b41-9c28-77ddf3f785b5Cited by top-tier papers21
- Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware MinimizationZiqing Fan, Shengchao Hu, Jiangchao Yao, Gang Niu et al.ICML 2024 · 35 citations
- Fundamental Convergence Analysis of Sharpness-Aware MinimizationPham Duy Khanh, Hoang-Chau Luong, Boris S. Mordukhovich, Dat Ba TranNeurIPS 2024 · 27 citations
- Momentum-SAM: Sharpness Aware Minimization without Computational OverheadMarlon Becker, Frederick Altrock, Benjamin RisseNeurIPS 2025 · 16 citations
- Implicit Regularization of Sharpness-Aware Minimization for Scale-Invariant ProblemsBingcong Li, Liang Zhang, Niao HeNeurIPS 2024 · 14 citations
- Efficient Sharpness-Aware Minimization for Molecular Graph Transformer ModelsYili Wang, Kaixiong Zhou, Ninghao Liu, Ying Wang et al.ICLR 2024 · 13 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 388 citations
Related papers
- Sharpness-Aware Training for FreeJiawei Du, Daquan Zhou, Jiashi Feng, Vincent Y. F. Tan et al.NeurIPS 2022 · 132 citations
- Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization TermYun Yue, Jiadi Jiang, Zhiling Ye, Ning Gao et al.KDD 2023 · 7 citations
- Why Does Sharpness-Aware Minimization Generalize Better Than SGD?Zixiang Chen, Junkai Zhang, Yiwen Kou, Xiangning Chen et al.NeurIPS 2023 · 32 citations
- Efficient Sharpness-aware Minimization for Improved Training of Neural NetworksJiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou et al.ICLR 2022 · 168 citations
- Revisiting Sharpness-Aware Minimization: A More Faithful and Effective ImplementationJianlong Chen, Zhiming ZhouICLR 2026 · 1 citation
