A Unified Stability Analysis of SAM vs SGD: Role of Data Coherence and Emergence of Simplicity Bias
Wei-Kai Chang, Rajiv Khanna
摘要
Understanding the dynamics of optimization in deep learning is increasingly important as models scale. While stochastic gradient descent (SGD) and its variants reliably find solutions that generalize well, the mechanisms driving this generalization remain unclear. Notably, these algorithms often prefer flatter or simpler minima-particularly in overparameterized settings. Prior work has linked flatness to generalization, and methods like Sharpness-Aware Minimization (SAM) explicitly encourage flatness, but a unified theory connecting data structure, optimization dynamics, and the nature of learned solutions is still lacking. In this work, we develop a linear stability framework that analyzes the behavior of SGD, random perturbations, and SAM-particularly in two-layer ReLU networks. Central to our analysis is a coherence measure that quantifies how gradient curvature aligns across data points, revealing why certain minima are stable and favored during training. (Code are available in: https://github.com/changwk1001/ Stability_Analysis_and_Simplicity-Bias.git)
Assuming w 0 ∼ N (0, I), we reduce to analyzing the quantity
, which captures the contraction or expansion behavior of the iterates under the sequence of update matrices. See more details discussion of assumption in appendix A.
The system is said to be linearly stable at w ⋆ under a given optimization method if the expected squared norm E[∥w k ∥ 2 ] remains bounded as k → ∞. A sufficient condition for this is that the spectral norm of the average update matrix E[ Ĵ⊤ t Ĵt ] is strictly less than 1. For full-batch gradient descent, this reduces to requiring η < 2/λ max (H).
More generally, in the presence of stochasticity and structure in the data, one can derive stability conditions involving both the Hessian spectrum and how curvature is distributed across examples. This motivates the use of a data-dependent coherence measure, which we introduce next.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Membership Privacy Risks of Sharpness Aware MinimizationYoung In Kim, Andrea Agiollo, Pratiksha Agrawal, Johannes O. Royset 等ICLR 2026 · 被引用 3 次
- Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware MinimizationChaewon Moon, Dongkuk Si, Chulhee YunICLR 2026 · 被引用 1 次
- Certification of Machine Learning Models via Directional SharpnessGefei Tan, Adrià Gascón, Sarah Meiklejohn, Mariana RaykovaUSENIX Security 2026
它引用的顶会 Paper20
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain 等NeurIPS 2020 · 被引用 503 次
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 被引用 385 次
相关 Paper
- A Precise Characterization of SGD Stability Using Loss Surface GeometryGregory Dexter, Borja Ocejo, S. Sathiya Keerthi, Aman Gupta 等ICLR 2024 · 被引用 2 次
- The alignment property of SGD noise and how it helps select flat minima: A stability analysisLei Wu, Mingze Wang, Weijie SuNeurIPS 2022 · 被引用 80 次
- The Implicit Regularization of Dynamical Stability in Stochastic Gradient DescentLei Wu, Weijie J. SuICML 2023 · 被引用 41 次
- Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering PhenomenonTongtong Liang, Dan Qiao, Yu-Xiang Wang, Rahul ParhiNeurIPS 2025 · 被引用 8 次
- On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGDTongcheng Zhang, Zhanpeng Zhou, Mingze Wang, Andi Han 等AAAI 2026
