Sharpness-Aware Training for Free
Jiawei Du, Daquan Zhou, Jiashi Feng, Vincent Y. F. Tan, Joey Tianyi Zhou
摘要
Modern deep neural networks (DNNs) have achieved state-of-the-art performances but are typically over-parameterized. The over-parameterization may result in undesirably large generalization error in the absence of other customized training strategies. Recently, a line of research under the name of Sharpness-Aware Minimization (SAM) has shown that minimizing a sharpness measure, which reflects the geometry of the loss landscape, can significantly reduce the generalization error. However, SAM-like methods incur a two-fold computational overhead of the given base optimizer (e.g. SGD) for approximating the sharpness measure. In this paper, we propose Sharpness-Aware Training for Free, or SAF, which mitigates the sharp landscape at almost zero additional computational cost over the base optimizer. Intuitively, SAF achieves this by avoiding sudden drops in the loss in the sharp local minima throughout the trajectory of the updates of the weights. Specifically, we suggest a novel trajectory loss, based on the KL-divergence between the outputs of DNNs with the current weights and past weights, as a replacement of the SAM's sharpness measure. This loss captures the rate of change of the training loss along the model's update trajectory. By minimizing it, SAF ensures the convergence to a flat minimum with improved generalization capabilities. Extensive empirical results show that SAF minimizes the sharpness in the same way that SAM does, yielding better results on the ImageNet dataset with essentially the same computational cost as the base optimizer. Our codes are available at https://github.com/AngusDujw/SAF .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- Demystify Mamba in Vision: A Linear Attention PerspectiveDongchen Han, Ziyi Wang, Zhuofan Xia, Yizeng Han 等NeurIPS 2024 · 被引用 287 次
- FasterViT: Fast Vision Transformers with Hierarchical AttentionAli Hatamizadeh, Greg Heinrich, Hongxu Yin, Andrew Tao 等ICLR 2024 · 被引用 132 次
- A Modern Look at the Relationship between Sharpness and GeneralizationMaksym Andriushchenko, Francesco Croce, Maximilian Müller, Matthias Hein 等ICML 2023 · 被引用 92 次
- FlatMatch: Bridging Labeled Data and Unlabeled Data with Cross-Sharpness for Semi-Supervised LearningZhuo Huang, Li Shen, Jun Yu, Bo Han 等NeurIPS 2023 · 被引用 50 次
- Make Continual Learning Stronger via C-FlatAng Bian, Wei Li, Hangjie Yuan, Chengrong Yu 等NeurIPS 2024 · 被引用 48 次
它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
相关 Paper
- Efficient Sharpness-aware Minimization for Improved Training of Neural NetworksJiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou 等ICLR 2022 · 被引用 168 次
- Make Sharpness-Aware Minimization Stronger: A Sparsified Perturbation ApproachPeng Mi, Li Shen, Tianhe Ren, Yiyi Zhou 等NeurIPS 2022 · 被引用 102 次
- Momentum-SAM: Sharpness Aware Minimization without Computational OverheadMarlon Becker, Frederick Altrock, Benjamin RisseNeurIPS 2025 · 被引用 16 次
- Revisiting Sharpness-Aware Minimization: A More Faithful and Effective ImplementationJianlong Chen, Zhiming ZhouICLR 2026 · 被引用 1 次
- Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization TermYun Yue, Jiadi Jiang, Zhiling Ye, Ning Gao 等KDD 2023 · 被引用 7 次
