Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity
Scott Pesme, Loucas Pillaud-Vivien, Nicolas Flammarion
摘要
Understanding the implicit bias of training algorithms is of crucial importance in order to explain the success of overparametrised neural networks. In this paper, we study the dynamics of stochastic gradient descent over diagonal linear networks through its continuous time version, namely stochastic gradient flow. We explicitly characterise the solution chosen by the stochastic flow and prove that it always enjoys better generalisation properties than that of gradient flow. Quite surprisingly, we show that the convergence speed of the training loss controls the magnitude of the biasing effect: the slower the convergence, the better the bias. To fully complete our analysis, we provide convergence guarantees for the dynamics. We also give experimental results which support our theoretical claims. Our findings highlight the fact that structured noise can induce better generalisation and they help explain the greater performances observed in practice of stochastic gradient descent over gradient descent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper65
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 被引用 190 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 被引用 92 次
- A Modern Look at the Relationship between Sharpness and GeneralizationMaksym Andriushchenko, Francesco Croce, Maximilian Müller, Matthias Hein 等ICML 2023 · 被引用 92 次
- The alignment property of SGD noise and how it helps select flat minima: A stability analysisLei Wu, Mingze Wang, Weijie SuNeurIPS 2022 · 被引用 80 次
它引用的顶会 Paper7
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth 等ICML 2021 · 被引用 85 次
- The Implicit Regularization of Stochastic Gradient Flow for Least SquaresAlnur Ali, Edgar Dobriban, Ryan J. TibshiraniICML 2020 · 被引用 83 次
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 52 次
- Stochastic Gradient and Langevin ProcessesXiang Cheng, Dong Yin, Peter L. Bartlett, Michael I. JordanICML 2020 · 被引用 51 次
相关 Paper
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 被引用 53 次
- Implicit Bias of (Stochastic) Gradient Descent for Rank-1 Linear Neural NetworkBochen Lyu, Zhanxing ZhuNeurIPS 2023 · 被引用 5 次
- Implicit Regularization for Group SparsityJiangyuan Li, Thanh Van Nguyen, Chinmay Hegde, Raymond K. W. WongICLR 2023 · 被引用 2 次
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
- SGD with Large Step Sizes Learns Sparse FeaturesMaksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionICML 2023 · 被引用 77 次
