Dissecting the Effects of SGD Noise in Distinct Regimes of Deep Learning
Antonio Sclocchi, Mario Geiger, Matthieu Wyart
摘要
Understanding when the noise in stochastic gradient descent (SGD) affects generalization of deep neural networks remains a challenge, complicated by the fact that networks can operate in distinct training regimes. Here we study how the magnitude of this noise affects performance as the size of the training set and the scale of initialization are varied. For gradient descent, is a key parameter that controls if the network is lazy'($\alpha\gg1$) or instead learns features ($\alpha\ll1$). For classification of MNIST and CIFAR10 images, our central results are: (i) obtaining phase diagrams for performance in the $(\alpha,T)$ plane. They show that SGD noise can be detrimental or instead useful depending on the training regime. Moreover, although increasing $T$ or decreasing $\alpha$ both allow the net to escape the lazy regime, these changes can have opposite effects on performance. (ii) Most importantly, we find that the characteristic temperature $T_c$ where the noise of SGD starts affecting the trained model (and eventually performance) is a power law of $P$. We relate this finding with the observation that key dynamical quantities, such as the total variation of weights during training, depend on both $T$ and $P$ as power laws. These results indicate that a key effect of SGD noise occurs late in training by affecting the stopping process whereby all data are fitted. Indeed, we argue that due to SGD noise, nets must develop a stronger signal', i.e. larger informative weights, to fit the data, leading to a longer training time. A stronger signal and a longer training time are also required when the size of the training set increases. We confirm these views in the perceptron model, where signal and noise can be precisely measured. Interestingly, exponents characterizing the effect of SGD depend on the density of data near the decision boundary, as we explain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MDN: Parallelizing Stepwise Momentum for Delta Linear AttentionYulong Huang, Xiang Liu, Hongxiang Huang, Xiaopeng LIN 等ICML 2026 · 被引用 1 次
- Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in GeneralizationTaesun Yeom, Taehyeok Ha, Jaeho LeeICML 2026
- A Risk Decomposition Framework for Pre-hoc Fine-tuning PredictionYuxiang Luo, Chen Wang, Nan TangICML 2026
- The Optimization Landscape of SGD Across the Feature Learning StrengthAlexander B. Atanasov, Alexandru Meterez, James B. Simon, Cengiz PehlevanICLR 2025
它引用的顶会 Paper4
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of StochasticityScott Pesme, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 135 次
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
- Failure and success of the spectral bias prediction for Laplace Kernel Ridge Regression: the case of low-dimensional dataUmberto M. Tomasini, Antonio Sclocchi, Matthieu WyartICML 2022 · 被引用 14 次
相关 Paper
- On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGDTongcheng Zhang, Zhanpeng Zhou, Mingze Wang, Andi Han 等AAAI 2026
- Benign Oscillation of Stochastic Gradient Descent with Large Learning RateMiao Lu, Beining Wu, Xiaodong Yang, Difan ZouICLR 2024 · 被引用 9 次
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 被引用 44 次
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
- Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler SubnetworksFeng Chen, Daniel Kunin, Atsushi Yamamura, Surya GanguliNeurIPS 2023 · 被引用 52 次
