The Implicit Regularization of Dynamical Stability in Stochastic Gradient Descent
Lei Wu, Weijie J. Su
摘要
In this paper, we study the implicit regularization of stochastic gradient descent (SGD) through the lens of dynamical stability (Wu et al., 2018). We start by revising existing stability analyses of SGD, showing how the Frobenius norm and trace of Hessian relate to different notions of stability. Notably, if a global minimum is linearly stable for SGD, then the trace of Hessian must be less than or equal to , where denotes the learning rate. By contrast, for gradient descent (GD), the stability imposes a similar constraint but only on the largest eigenvalue of Hessian. We then turn to analyze the generalization properties of these stable minima, focusing specifically on two-layer ReLU networks and diagonal linear networks. Notably, we establish the equivalence between these metrics of sharpness and certain parameter norms for the two models, which allows us to show that the stable minima of SGD provably generalize well. By contrast, the stability-induced regularization of GD is provably too weak to ensure satisfactory generalization. This discrepancy provides an explanation of why SGD often generalizes better than GD. Note that the learning rate (LR) plays a pivotal role in the strength of stability-induced regularization. As the LR increases, the regularization effect becomes more pronounced, elucidating why SGD with a larger LR consistently demonstrates superior generalization capabilities. Additionally, numerical experiments are provided to support our theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- On the Generalization Properties of Diffusion ModelsPuheng Li, Zhong Li, Huishuai Zhang, Jiang BianNeurIPS 2023 · 被引用 86 次
- Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better GeneralizationKaiyue Wen, Zhiyuan Li, Tengyu MaNeurIPS 2023 · 被引用 53 次
- Towards Theoretical Understandings of Self-Consuming Generative ModelsShi Fu, Sen Zhang, Yingjie Wang, Xinmei Tian 等ICML 2024 · 被引用 26 次
- Improving Generalization and Convergence by Enhancing Implicit RegularizationMingze Wang, Jinbo Wang, Haotian He, Zilin Wang 等NeurIPS 2024 · 被引用 21 次
- Decentralized SGD and Average-direction SAM are Asymptotically EquivalentTongtian Zhu, Fengxiang He, Kaixuan Chen, Mingli Song 等ICML 2023 · 被引用 21 次
它引用的顶会 Paper15
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 被引用 235 次
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 被引用 235 次
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 被引用 155 次
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of StochasticityScott Pesme, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 135 次
相关 Paper
- The alignment property of SGD noise and how it helps select flat minima: A stability analysisLei Wu, Mingze Wang, Weijie SuNeurIPS 2022 · 被引用 80 次
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 被引用 44 次
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 被引用 21 次
- Conflicting Biases at the Edge of Stability: Norm versus Sharpness RegularizationMaria Matveev, Vit Fojtik, Hung-Hsu Chou, Gitta Kutyniok 等ICML 2026
- A Precise Characterization of SGD Stability Using Loss Surface GeometryGregory Dexter, Borja Ocejo, S. Sathiya Keerthi, Aman Gupta 等ICLR 2024 · 被引用 2 次
