Strength of Minibatch Noise in SGD
Liu Ziyin, Kangqiao Liu, Takashi Mori, Masahito Ueda
摘要
The noise in stochastic gradient descent (SGD), caused by minibatch sampling, is poorly understood despite its practical importance in deep learning. This work presents the first systematic study of the SGD noise and fluctuations close to a local minimum. We first analyze the SGD noise in linear regression in detail and then derive a general formula for approximating SGD noise in different types of minima. For application, our results (1) provide insight into the stability of training a neural network, (2) suggest that a large learning rate can help generalization by introducing an implicit regularization, (3) explain why the linear learning rate-batchsize scaling law fails at a large learning rate or at a small batchsize and (4) can provide an understanding of how discrete-time nature of SGD affects the recently discovered power-law phenomenon of SGD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- The alignment property of SGD noise and how it helps select flat minima: A stability analysisLei Wu, Mingze Wang, Weijie SuNeurIPS 2022 · 被引用 80 次
- SGD with Large Step Sizes Learns Sparse FeaturesMaksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionICML 2023 · 被引用 77 次
- Parameter Symmetry and Noise Equilibrium of Stochastic Gradient DescentLiu Ziyin, Mingze Wang, Hongchao Li, Lei WuNeurIPS 2024 · 被引用 23 次
- Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High DimensionsCourtney Paquette, Elliot Paquette, Ben Adlam, Jeffrey PenningtonNeurIPS 2022 · 被引用 22 次
- Decentralized SGD and Average-direction SAM are Asymptotically EquivalentTongtian Zhu, Fengxiang He, Kaixuan Chen, Mingli Song 等ICML 2023 · 被引用 21 次
它引用的顶会 Paper5
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 被引用 165 次
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan 等ICML 2020 · 被引用 125 次
- Multiplicative Noise and Heavy Tails in Stochastic OptimizationLiam Hodgkinson, Michael W. MahoneyICML 2021 · 被引用 90 次
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient DescentKangqiao Liu, Liu Ziyin, Masahito UedaICML 2021 · 被引用 46 次
- SGD Can Converge to Local MaximaLiu Ziyin, Botao Li, James B. Simon, Masahito UedaICLR 2022 · 被引用 18 次
相关 Paper
- The Implicit Regularization of Dynamical Stability in Stochastic Gradient DescentLei Wu, Weijie J. SuICML 2023 · 被引用 41 次
- Benign Oscillation of Stochastic Gradient Descent with Large Learning RateMiao Lu, Beining Wu, Xiaodong Yang, Difan ZouICLR 2024 · 被引用 9 次
- How much does Initialization Affect Generalization?Sameera Ramasinghe, Lachlan Ewen MacDonald, Moshiur R. Farazi, Hemanth Saratchandran 等ICML 2023 · 被引用 9 次
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
