Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High Dimensions
Courtney Paquette, Elliot Paquette, Ben Adlam, Jeffrey Pennington
摘要
Stochastic gradient descent (SGD) is a pillar of modern machine learning, serving as the go-to optimization algorithm for a diverse array of problems. While the empirical success of SGD is often attributed to its computational efficiency and favorable generalization behavior, neither effect is well understood and disentangling them remains an open problem. Even in the simple setting of convex quadratic problems, worst-case analyses give an asymptotic convergence rate for SGD that is no better than full-batch gradient descent (GD), and the purported implicit regularization effects of SGD lack a precise explanation. In this work, we study the dynamics of multi-pass SGD on high-dimensional convex quadratics and establish an asymptotic equivalence to a stochastic differential equation, which we call homogenized stochastic gradient descent (HSGD), whose solutions we characterize explicitly in terms of a Volterra integral equation. These results yield precise formulas for the learning and risk trajectories, which reveal a mechanism of implicit conditioning that explains the efficiency of SGD relative to GD. We also prove that the noise from SGD negatively impacts generalization performance, ruling out the possibility of any type of implicit regularization in this context. Finally, we show how to adapt the HSGD formalism to include streaming SGD, which allows us to produce an exact prediction for the excess risk of multi-pass SGD relative to that of streaming SGD (bootstrap risk).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SAM operates far from home: eigenvalue regularization as a dynamical phenomenonAtish Agarwala, Yann N. DauphinICML 2023 · 被引用 26 次
- Beyond Implicit Bias: The Insignificance of SGD Noise in Online LearningNikhil Vyas, Depen Morwani, Rosie Zhao, Gal Kaplun 等ICML 2024 · 被引用 8 次
- Emergence of heavy tails in homogenized stochastic gradient descentZhezhe Jiao, Martin Keller-ResselNeurIPS 2024 · 被引用 6 次
- The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate AlgorithmsElizabeth Collins-Woodfin, Inbar Seroussi, Begoña García Malaxechebarría, Andrew W. Mackenzie 等NeurIPS 2024 · 被引用 3 次
- Statistical Guarantees for High-Dimensional Stochastic Gradient DescentJiaqi Li, Zhipeng Lou, Johannes Schmidt-Hieber, Wei Biao WuNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper13
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 被引用 235 次
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 被引用 165 次
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 被引用 133 次
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 被引用 111 次
相关 Paper
- Risk Bounds of Multi-Pass SGD for Least Squares in the Interpolation RegimeDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu 等NeurIPS 2022 · 被引用 9 次
- The Benefits of Implicit Regularization from SGD in Least Squares ProblemsDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu 等NeurIPS 2021 · 被引用 41 次
- Trajectory of Mini-Batch Momentum: Batch Size Saturation and Convergence in High DimensionsKiwon Lee, Andrew N. Cheng, Elliot Paquette, Courtney PaquetteNeurIPS 2022 · 被引用 22 次
- Can Implicit Bias Explain Generalization? Stochastic Convex Optimization as a Case StudyAssaf Dauber, Meir Feder, Tomer Koren, Roi LivniNeurIPS 2020 · 被引用 26 次
- SGD: The Role of Implicit Regularization, Batch-size and Multiple-epochsAyush Sekhari, Karthik Sridharan, Satyen KaleNeurIPS 2021 · 被引用 36 次
