Sharpness-aware Minimization for Efficiently Improving Generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam Neyshabur
摘要
In today's heavily overparameterized models, the value of the training loss provides few guarantees on model generalization ability. Indeed, optimizing only the training loss value, as is commonly done, can easily lead to suboptimal model quality. Motivated by prior work connecting the geometry of the loss landscape and generalization, we introduce a novel, effective procedure for instead simultaneously minimizing loss value and loss sharpness. In particular, our procedure, Sharpness-Aware Minimization (SAM), seeks parameters that lie in neighborhoods having uniformly low loss; this formulation results in a minmax optimization problem on which gradient descent can be performed efficiently. We present empirical results showing that SAM improves model generalization across a variety of benchmark datasets (e.g., CIFAR-10, 100, Ima-geNet, finetuning tasks) and models, yielding novel state-of-the-art performance for several. Additionally, we find that SAM natively provides robustness to label noise on par with that provided by state-of-the-art procedures that specifically target learning with noisy labels. We open source our code at https: //github.com/google-research/sam . * Work done as part of the Google AI Residency program.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper650
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 被引用 653 次
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho 等NeurIPS 2021 · 被引用 630 次
它引用的顶会 Paper5
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- O2U-Net: A Simple Noisy Label Detection Approach for Deep Neural NetworksJinchi Huang, Lie Qu, Rongfei Jia, Binqiang ZhaoICCV 2019 · 被引用 276 次
- Improved Sample Complexities for Deep Neural Networks and Robust Classification via an All-Layer MarginColin Wei, Tengyu MaICLR 2020 · 被引用 91 次
- The intriguing role of module criticality in the generalization of deep networksNiladri S. Chatterji, Behnam Neyshabur, Hanie SedghiICLR 2020 · 被引用 59 次
- Exploring the Vulnerability of Deep Neural Networks: A Study of Parameter CorruptionXu Sun, Zhiyuan Zhang, Xuancheng Ren, Ruixuan Luo 等AAAI 2021 · 被引用 45 次
相关 Paper
- Revisiting Sharpness-Aware Minimization: A More Faithful and Effective ImplementationJianlong Chen, Zhiming ZhouICLR 2026 · 被引用 1 次
- Why Does Sharpness-Aware Minimization Generalize Better Than SGD?Zixiang Chen, Junkai Zhang, Yiwen Kou, Xiangning Chen 等NeurIPS 2023 · 被引用 32 次
- Random Sharpness-Aware MinimizationYong Liu, Siqi Mai, Minhao Cheng, Xiangning Chen 等NeurIPS 2022 · 被引用 38 次
- Enhancing Sharpness-Aware Optimization Through Variance SuppressionBingcong Li, Georgios B. GiannakisNeurIPS 2023 · 被引用 47 次
- Align-SAM: Seeking Flatter Minima for Better Cross-Subset AlignmentVan-Anh Nguyen, Mehrtash Harandi, Thanh-Toan Do, Linh Ngo Van 等ICLR 2026
