On the Origin of Implicit Regularization in Stochastic Gradient Descent
Samuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham De
摘要
For infinitesimal learning rates, stochastic gradient descent (SGD) follows the path of gradient flow on the full batch loss function. However moderately large learning rates can achieve higher test accuracies, and this generalization benefit is not explained by convergence bounds, since the learning rate which maximizes test accuracy is often larger than the learning rate which minimizes training loss. To interpret this phenomenon we prove that for SGD with random shuffling, the mean SGD iterate also stays close to the path of gradient flow if the learning rate is small and finite, but on a modified loss. This modified loss is composed of the original loss function and an implicit regularizer, which penalizes the norms of the minibatch gradients. Under mild assumptions, when the batch size is small the scale of the implicit regularization term is proportional to the ratio of the learning rate to the batch size. We verify empirically that explicitly including the implicit regularizer in the loss can enhance the test accuracy when the learning rate is small. INTRODUCTION In the limit of vanishing learning rates, stochastic gradient descent with minibatch gradients (SGD) follows the path of gradient flow on the full batch loss function (Yaida, 2019) . However in deep networks, SGD often achieves higher test accuracies when the learning rate is moderately large (LeCun et al., 2012; Keskar et al., 2017) . This generalization benefit is not explained by convergence rate bounds (Ma et al., 2018; Zhang et al., 2019) , because it arises even for large compute budgets for which smaller learning rates often achieve lower training losses (Smith et al., 2020). Although many authors have studied this phenomenon (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper84
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren 等NeurIPS 2025 · 被引用 314 次
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 被引用 204 次
- Manipulating SGD with Data Ordering AttacksIlia Shumailov, Zakhar Shumaylov, Dmitry Kazhdan, Yiren Zhao 等NeurIPS 2021 · 被引用 125 次
- On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)Zhiyuan Li, Sadhika Malladi, Sanjeev AroraNeurIPS 2021 · 被引用 107 次
- A Modern Look at the Relationship between Sharpness and GeneralizationMaksym Andriushchenko, Francesco Croce, Maximilian Müller, Matthias Hein 等ICML 2023 · 被引用 92 次
它引用的顶会 Paper4
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 被引用 173 次
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang 等ICML 2020 · 被引用 122 次
相关 Paper
- Implicit Regularization of SGD Reduces Shortcut LearningNahal Mirzaie, Alireza Alipanah, Ali Abbasi, Amirmahdi Farzane 等ICLR 2026
- Implicit regularization in Heavy-ball momentum accelerated stochastic gradient descentAvrajit Ghosh, He Lyu, Xitong Zhang, Rongrong WangICLR 2023 · 被引用 1 次
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 被引用 44 次
- SGD: The Role of Implicit Regularization, Batch-size and Multiple-epochsAyush Sekhari, Karthik Sridharan, Satyen KaleNeurIPS 2021 · 被引用 36 次
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 被引用 155 次
