GA-SAM: Gradient-Strength based Adaptive Sharpness-Aware Minimization for Improved Generalization
Zhiyuan Zhang, Ruixuan Luo, Qi Su, Xu Sun
摘要
Recently, Sharpness-Aware Minimization (SAM) algorithm has shown state-of-the-art generalization abilities in vision tasks. It demonstrates that flat minima tend to imply better generalization abilities. However, it has some difficulty implying SAM to some natural language tasks, especially to models with drastic gradient changes, such as RNNs. In this work, we analyze the relation between the flatness of the local minimum and its generalization ability from a novel and straightforward theoretical perspective. We propose that the shift of the training and test distributions can be equivalently seen as a virtual parameter corruption or perturbation, which can explain why flat minima that are robust against parameter corruptions or perturbations have better generalization performances. On its basis, we propose a Gradient-Strength based Adaptive Sharpness-Aware Minimization (GA-SAM) algorithm to help to learn algorithms find flat minima that generalize better. Results in various language benchmarks validate the effectiveness of the proposed GA-SAM algorithm on natural language tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- FlatMatch: Bridging Labeled Data and Unlabeled Data with Cross-Sharpness for Semi-Supervised LearningZhuo Huang, Li Shen, Jun Yu, Bo Han 等NeurIPS 2023 · 被引用 50 次
- Enhancing Sharpness-Aware Optimization Through Variance SuppressionBingcong Li, Georgios B. GiannakisNeurIPS 2023 · 被引用 47 次
- Fed-FA: Theoretically Modeling Client Data Divergence for Federated Language Backdoor DefenseZhiyuan Zhang, Deli Chen, Hao Zhou, Fandong Meng 等NeurIPS 2023 · 被引用 10 次
- Robust Generalization Against Photon-Limited Corruptions via Worst-Case Sharpness MinimizationZhuo Huang, Miaoxi Zhu, Xiaobo Xia, Li Shen 等CVPR 2023
- Beyond Sharpness: The Role of Nonuniformity in GeneralizationYingcong Zhou, Pingfan Wu, Li Wang, Zhiguo Fu 等AAAI 2026
它引用的顶会 Paper11
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 被引用 622 次
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 被引用 385 次
相关 Paper
- Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Hao Zou 等CVPR 2023
- When Do Flat Minima Optimizers Work?Jean Kaddour, Linqing Liu, Ricardo Silva, Matt J. KusnerNeurIPS 2022 · 被引用 102 次
- Tilted Sharpness-Aware MinimizationTian Li, Tianyi Zhou, Jeff A. BilmesICML 2025
- Align-SAM: Seeking Flatter Minima for Better Cross-Subset AlignmentVan-Anh Nguyen, Mehrtash Harandi, Thanh-Toan Do, Linh Ngo Van 等ICLR 2026
- Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late In TrainingZhanpeng Zhou, Mingze Wang, Yuchen Mao, Bingrui Li 等ICLR 2025
