On Saddle Point Avoidance and Stationary Distribution of Sharpness-Aware Minimization
Tao Sun, Fan Jia, Bao Wang
摘要
We revisit the Sharpness-Aware Minimization (SAM) method and its unnormalized variant (USAM), which are widely used in training neural networks. In this work, we introduce three new theoretical results to enhance the understanding of (U)SAM. Firstly, we prove that deterministic USAM always converges to a local minimum and exhibits local acceleration compared to standard gradient descent, offering theoretical guarantees for its superior performance near minima. Secondly, we analyze the trade-off between convergence and escaping saddle points, showing that while SAM's behavior may prevent convergence in some cases, it effectively helps escape saddle points, leading to better overall optimization. Lastly, we demonstrate that USAM behaves like a Markov chain in the presence of stationary noise, with the objective values on the stationary distribution being lower than those achieved by SGD, highlighting USAM's advantages in noisy settings. All of our theoretical results are supported by the numerical experiments reported in this paper.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Stability Analysis of Sharpness-Aware MinimizationHoki Kim, Jinseong Park, Yujin Choi, Jaewook LeeICML 2026 · 被引用 18 次
- Sharpness-Aware Minimization Can Hallucinate MinimizersChanwoong Park, Uijeong Jang, Ernest Ryu, Insoon YangICML 2026
- An SDE for Modeling SAM: Theory and InsightsEnea Monzio Compagnoni, Luca Biggio, Antonio Orvieto, Frank Norbert Proske 等ICML 2023 · 被引用 25 次
- Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late In TrainingZhanpeng Zhou, Mingze Wang, Yuchen Mao, Bingrui Li 等ICLR 2025
- Why Does Sharpness-Aware Minimization Generalize Better Than SGD?Zixiang Chen, Junkai Zhang, Yiwen Kou, Xiangning Chen 等NeurIPS 2023 · 被引用 32 次
