Lune

KDD2026顶会

On Saddle Point Avoidance and Stationary Distribution of Sharpness-Aware Minimization

Tao Sun, Fan Jia, Bao Wang

2026年份
1被引次数

摘要

We revisit the Sharpness-Aware Minimization (SAM) method and its unnormalized variant (USAM), which are widely used in training neural networks. In this work, we introduce three new theoretical results to enhance the understanding of (U)SAM. Firstly, we prove that deterministic USAM always converges to a local minimum and exhibits local acceleration compared to standard gradient descent, offering theoretical guarantees for its superior performance near minima. Secondly, we analyze the trade-off between convergence and escaping saddle points, showing that while SAM's behavior may prevent convergence in some cases, it effectively helps escape saddle points, leading to better overall optimization. Lastly, we demonstrate that USAM behaves like a Markov chain in the presence of stationary noise, with the objective values on the stationary distribution being lower than those achieved by SGD, highlighting USAM's advantages in noisy settings. All of our theoretical results are supported by the numerical experiments reported in this paper.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖