On the Noisy Gradient Descent that Generalizes as SGD
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, Zhanxing Zhu
摘要
The gradient noise of SGD is considered to play a central role in the observed strong generalization abilities of deep learning. While past studies confirm that the magnitude and covariance structure of gradient noise are critical for regularization, it remains unclear whether or not the class of noise distributions is important. In this work we provide negative results by showing that noises in classes different from the SGD noise can also effectively regularize gradient descent. Our finding is based on a novel observation on the structure of the SGD noise: it is the multiplication of the gradient matrix and a sampling noise that arises from the mini-batch sampling procedure. Moreover, the sampling noises unify two kinds of gradient regularizing noises that belong to the Gaussian class: the one using (scaled) Fisher as covariance and the one using the gradient covariance of SGD as covariance. Finally, thanks to the flexibility of choosing noise class, an algorithm is proposed to perform noisy gradient descent that generalizes well, the variant of which even benefits large batch SGD training without hurting generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang 等CVPR 2022 · 被引用 264 次
- What Happens after SGD Reaches Zero Loss? --A Mathematical FrameworkZhiyuan Li, Tianhao Wang, Sanjeev AroraICLR 2022 · 被引用 121 次
- Open-set Label Noise Can Improve Robustness Against Inherent Label NoiseHongxin Wei, Lue Tao, Renchunzi Xie, Bo AnNeurIPS 2021 · 被引用 113 次
- On Training Implicit ModelsZhengyang Geng, Xin-Yu Zhang, Shaojie Bai, Yisen Wang 等NeurIPS 2021 · 被引用 111 次
- On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)Zhiyuan Li, Sadhika Malladi, Sanjeev AroraNeurIPS 2021 · 被引用 107 次
相关 Paper
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 被引用 44 次
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 被引用 155 次
- Augment Your Batch: Improving Generalization Through Instance RepetitionElad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi 等CVPR 2020
- Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to Improve GeneralizationZeke Xie, Li Yuan, Zhanxing Zhu, Masashi SugiyamaICML 2021 · 被引用 39 次
