Stability Analysis of Sharpness-Aware Minimization
Hoki Kim, Jinseong Park, Yujin Choi, Jaewook Lee
摘要
Sharpness-aware minimization (SAM) is a training method that seeks to find flat minima in deep learning, resulting in state-of-the-art performance across various domains. Instead of minimizing the loss of the current weights, SAM minimizes the worst-case loss in its neighborhood in the parameter space. In this paper, we investigate the convergence instability of SAM near a saddle point. Using the qualitative theory of dynamical systems, we explain how SAM becomes stuck in the saddle point and theoretically prove that the saddle point can become an attractor under SAM dynamics. Additionally, we show that this convergence instability can also occur in stochastic dynamical systems by establishing the diffusion of SAM. We prove that SAM diffusion is worse than that of vanilla gradient descent in terms of saddle point escape. Finally, we demonstrate that often overlooked training tricks, momentum and batch-size, might be important to mitigate the convergence instability and achieve high generalization performance. Our theoretical and empirical results are thoroughly verified through experiments on several well-known optimization problems and benchmark tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- The Crucial Role of Normalization in Sharpness-Aware MinimizationYan Dai, Kwangjun Ahn, Suvrit SraNeurIPS 2023 · 被引用 36 次
- Practical Sharpness-Aware Minimization Cannot Converge All the Way to OptimaDongkuk Si, Chulhee YunNeurIPS 2023 · 被引用 34 次
- An SDE for Modeling SAM: Theory and InsightsEnea Monzio Compagnoni, Luca Biggio, Antonio Orvieto, Frank Norbert Proske 等ICML 2023 · 被引用 25 次
- Decentralized SGD and Average-direction SAM are Asymptotically EquivalentTongtian Zhu, Fengxiang He, Kaixuan Chen, Mingli Song 等ICML 2023 · 被引用 21 次
- Differentially Private Sharpness-Aware TrainingJinseong Park, Hoki Kim, Yujin Choi, Jaewook LeeICML 2023 · 被引用 15 次
它引用的顶会 Paper14
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 被引用 388 次
- Surrogate Gap Minimization Improves Sharpness-Aware TrainingJuntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui 等ICLR 2022 · 被引用 213 次
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 被引用 190 次
相关 Paper
- On Saddle Point Avoidance and Stationary Distribution of Sharpness-Aware MinimizationTao Sun, Fan Jia, Bao WangKDD 2026 · 被引用 1 次
- Improving Sharpness-Aware Minimization by LookaheadRunsheng Yu, Youzhi Zhang, James T. KwokICML 2024 · 被引用 1 次
- Sharpness-Aware Minimization Can Hallucinate MinimizersChanwoong Park, Uijeong Jang, Ernest Ryu, Insoon YangICML 2026
- Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late In TrainingZhanpeng Zhou, Mingze Wang, Yuchen Mao, Bingrui Li 等ICLR 2025
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 被引用 3 次
