On Saddle Point Avoidance and Stationary Distribution of Sharpness-Aware Minimization
Tao Sun, Fan Jia, Bao Wang
Abstract
We revisit the Sharpness-Aware Minimization (SAM) method and its unnormalized variant (USAM), which are widely used in training neural networks. In this work, we introduce three new theoretical results to enhance the understanding of (U)SAM. Firstly, we prove that deterministic USAM always converges to a local minimum and exhibits local acceleration compared to standard gradient descent, offering theoretical guarantees for its superior performance near minima. Secondly, we analyze the trade-off between convergence and escaping saddle points, showing that while SAM's behavior may prevent convergence in some cases, it effectively helps escape saddle points, leading to better overall optimization. Lastly, we demonstrate that USAM behaves like a Markov chain in the presence of stationary noise, with the objective values on the stationary distribution being lower than those achieved by SGD, highlighting USAM's advantages in noisy settings. All of our theoretical results are supported by the numerical experiments reported in this paper.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a9f7484a-59af-4915-8539-62c22ae3eb3fRelated papers
- Stability Analysis of Sharpness-Aware MinimizationHoki Kim, Jinseong Park, Yujin Choi, Jaewook LeeICML 2026 · 18 citations
- Sharpness-Aware Minimization Can Hallucinate MinimizersChanwoong Park, Uijeong Jang, Ernest Ryu, Insoon YangICML 2026
- An SDE for Modeling SAM: Theory and InsightsEnea Monzio Compagnoni, Luca Biggio, Antonio Orvieto, Frank Norbert Proske et al.ICML 2023 · 25 citations
- Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late In TrainingZhanpeng Zhou, Mingze Wang, Yuchen Mao, Bingrui Li et al.ICLR 2025
- Why Does Sharpness-Aware Minimization Generalize Better Than SGD?Zixiang Chen, Junkai Zhang, Yiwen Kou, Xiangning Chen et al.NeurIPS 2023 · 32 citations
