Enhancing Sharpness-Aware Optimization Through Variance Suppression
Bingcong Li, Georgios B. Giannakis
摘要
Sharpness-aware minimization (SAM) has well documented merits in enhancing generalization of deep neural networks, even without sizable data augmentation. Embracing the geometry of the loss function, where neighborhoods of 'flat minima' heighten generalization ability, SAM seeks 'flat valleys' by minimizing the maximum loss caused by an adversary perturbing parameters within the neighborhood. Although critical to account for sharpness of the loss function, such an 'over-friendly adversary' can curtail the outmost level of generalization. The novel approach of this contribution fosters stabilization of adversaries through variance suppression (VaSSO) to avoid such friendliness. VaSSO's provable stability safeguards its numerical improvement over SAM in model-agnostic tasks, including image classification and machine translation. In addition, experiments confirm that VaSSO endows SAM with robustness against high levels of label noise. Code is available at https://github.com/ BingcongLi/VaSSO . Introduction Despite deep neural networks (DNNs) have advanced the concept of "learning from data," and markedly improved performance across several applications in vision and language (Devlin et al., 2018; Tom et al., 2020) , their overparametrized nature renders the tendency to overfit on training data (Zhang et al., 2021a). This has led to concerns in generalization, which is a practically underscored perspective yet typically suffers from a gap relative to the training performance. Improving generalizability is challenging. Common approaches include (model) regularization and data augmentation (Srivastava et al., 2014) . While it is the default choice to integrate regularization such as weight decay and dropout into training, these methods are often insufficient for DNNs especially when coping with complicated network architectures (Chen et al., 2022) . Another line of effort resorts to suitable optimization schemes attempting to find a generalizable local minimum. For example, SGD is more preferable than Adam on certain overparameterized problems since it converges to maximum margin solutions (Wilson et al., 2017) . Decoupling weight decay from Adam also empirically facilitates generalizability (Loshchilov and Hutter, 2017) . Unfortunately, the underlying mechanism remains unveiled, and whether the generalization merits carry over to other intricate learning tasks calls for additional theoretical elaboration. Our main focus, sharpness aware minimization (SAM), is a highly compelling optimization approach that facilitates state-of-the-art generalizability by exploiting sharpness of loss landscape (Foret et al., 2021; Chen et al., 2022) . A high-level interpretation of sharpness is how violently the loss fluctuates within a neighborhood. It has been shown through large-scale empirical studies that sharpness-based measures highly
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware MinimizationZiqing Fan, Shengchao Hu, Jiangchao Yao, Gang Niu 等ICML 2024 · 被引用 35 次
- Fundamental Convergence Analysis of Sharpness-Aware MinimizationPham Duy Khanh, Hoang-Chau Luong, Boris S. Mordukhovich, Dat Ba TranNeurIPS 2024 · 被引用 27 次
- Momentum-SAM: Sharpness Aware Minimization without Computational OverheadMarlon Becker, Frederick Altrock, Benjamin RisseNeurIPS 2025 · 被引用 16 次
- Implicit Regularization of Sharpness-Aware Minimization for Scale-Invariant ProblemsBingcong Li, Liang Zhang, Niao HeNeurIPS 2024 · 被引用 14 次
- Efficient Sharpness-Aware Minimization for Molecular Graph Transformer ModelsYili Wang, Kaixiong Zhou, Ninghao Liu, Ying Wang 等ICLR 2024 · 被引用 13 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 被引用 388 次
相关 Paper
- Sharpness-Aware Training for FreeJiawei Du, Daquan Zhou, Jiashi Feng, Vincent Y. F. Tan 等NeurIPS 2022 · 被引用 132 次
- Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization TermYun Yue, Jiadi Jiang, Zhiling Ye, Ning Gao 等KDD 2023 · 被引用 7 次
- Why Does Sharpness-Aware Minimization Generalize Better Than SGD?Zixiang Chen, Junkai Zhang, Yiwen Kou, Xiangning Chen 等NeurIPS 2023 · 被引用 32 次
- Efficient Sharpness-aware Minimization for Improved Training of Neural NetworksJiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou 等ICLR 2022 · 被引用 168 次
- Revisiting Sharpness-Aware Minimization: A More Faithful and Effective ImplementationJianlong Chen, Zhiming ZhouICLR 2026 · 被引用 1 次
