A Unified Stability Analysis of SAM vs SGD: Role of Data Coherence and Emergence of Simplicity Bias
Wei-Kai Chang, Rajiv Khanna
Abstract
Understanding the dynamics of optimization in deep learning is increasingly important as models scale. While stochastic gradient descent (SGD) and its variants reliably find solutions that generalize well, the mechanisms driving this generalization remain unclear. Notably, these algorithms often prefer flatter or simpler minima-particularly in overparameterized settings. Prior work has linked flatness to generalization, and methods like Sharpness-Aware Minimization (SAM) explicitly encourage flatness, but a unified theory connecting data structure, optimization dynamics, and the nature of learned solutions is still lacking. In this work, we develop a linear stability framework that analyzes the behavior of SGD, random perturbations, and SAM-particularly in two-layer ReLU networks. Central to our analysis is a coherence measure that quantifies how gradient curvature aligns across data points, revealing why certain minima are stable and favored during training. (Code are available in: https://github.com/changwk1001/ Stability_Analysis_and_Simplicity-Bias.git)
Assuming w 0 ∼ N (0, I), we reduce to analyzing the quantity
, which captures the contraction or expansion behavior of the iterates under the sequence of update matrices. See more details discussion of assumption in appendix A.
The system is said to be linearly stable at w ⋆ under a given optimization method if the expected squared norm E[∥w k ∥ 2 ] remains bounded as k → ∞. A sufficient condition for this is that the spectral norm of the average update matrix E[ Ĵ⊤ t Ĵt ] is strictly less than 1. For full-batch gradient descent, this reduces to requiring η < 2/λ max (H).
More generally, in the presence of stochasticity and structure in the data, one can derive stability conditions involving both the Hessian spectrum and how curvature is distributed across examples. This motivates the use of a data-dependent coherence measure, which we introduce next.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f288f762-dbdd-40d1-8bd4-243dc4715497Cited by top-tier papers3
- Membership Privacy Risks of Sharpness Aware MinimizationYoung In Kim, Andrea Agiollo, Pratiksha Agrawal, Johannes O. Royset et al.ICLR 2026 · 3 citations
- Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware MinimizationChaewon Moon, Dongkuk Si, Chulhee YunICLR 2026 · 1 citation
- Certification of Machine Learning Models via Directional SharpnessGefei Tan, Adrià Gascón, Sarah Meiklejohn, Mariana RaykovaUSENIX Security 2026
Builds on20
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 385 citations
Related papers
- A Precise Characterization of SGD Stability Using Loss Surface GeometryGregory Dexter, Borja Ocejo, S. Sathiya Keerthi, Aman Gupta et al.ICLR 2024 · 2 citations
- The alignment property of SGD noise and how it helps select flat minima: A stability analysisLei Wu, Mingze Wang, Weijie SuNeurIPS 2022 · 80 citations
- The Implicit Regularization of Dynamical Stability in Stochastic Gradient DescentLei Wu, Weijie J. SuICML 2023 · 41 citations
- Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering PhenomenonTongtong Liang, Dan Qiao, Yu-Xiang Wang, Rahul ParhiNeurIPS 2025 · 8 citations
- On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGDTongcheng Zhang, Zhanpeng Zhou, Mingze Wang, Andi Han et al.AAAI 2026
