Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves Generalization
Xingxuan Zhang, Renzhe Xu, Han Yu, Hao Zou, Peng Cui
摘要
Recently, flat minima are proven to be effective for improving generalization and sharpness-aware minimization (SAM) achieves state-of-the-art performance. Yet the current definition of flatness discussed in SAM and its followups are limited to the zeroth-order flatness (i.e., the worstcase loss within a perturbation radius). We show that the zeroth-order flatness can be insufficient to discriminate minima with low generalization error from those with high generalization error both when there is a single minimum or multiple minima within the given perturbation radius. Thus we present first-order flatness, a stronger measure of flatness focusing on the maximal gradient norm within a perturbation radius which bounds both the maximal eigenvalue of Hessian at local minima and the regularization function of SAM. We also present a novel training procedure named Gradient norm Aware Minimization (GAM) to seek minima with uniformly small curvature across all directions. Experimental results show that GAM improves the generalization of models trained with current optimizers such as SGD and AdamW on various datasets and networks. Furthermore, we show that GAM can help SAM find flatter minima and achieve better generalization. The code is available at https://github.com/xxgege/GAM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- Sharpness-Aware Minimization Leads to Low-Rank FeaturesMaksym Andriushchenko, Dara Bahri, Hossein Mobahi, Nicolas FlammarionNeurIPS 2023 · 被引用 48 次
- Make Continual Learning Stronger via C-FlatAng Bian, Wei Li, Hangjie Yuan, Chengrong Yu 等NeurIPS 2024 · 被引用 48 次
- Flatness-Aware Minimization for Domain GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Yancheng Dong 等ICCV 2023 · 被引用 37 次
- Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware MinimizationZiqing Fan, Shengchao Hu, Jiangchao Yao, Gang Niu 等ICML 2024 · 被引用 35 次
- Practical Sharpness-Aware Minimization Cannot Converge All the Way to OptimaDongkuk Si, Chulhee YunNeurIPS 2023 · 被引用 34 次
它引用的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
相关 Paper
- Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization TermYun Yue, Jiadi Jiang, Zhiling Ye, Ning Gao 等KDD 2023 · 被引用 7 次
- Fix the Loss, Not the Radius: Rethinking the Adversarial Perturbation of Sharpness-Aware MinimizationJinping Wang, Qinhan Liu, Zhiwu Xie, Zhiqiang GaoICML 2026
- Revisiting Sharpness-Aware Minimization: A More Faithful and Effective ImplementationJianlong Chen, Zhiming ZhouICLR 2026 · 被引用 1 次
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 被引用 3 次
- Improving Sharpness-Aware Minimization by LookaheadRunsheng Yu, Youzhi Zhang, James T. KwokICML 2024 · 被引用 1 次
