SGD: The Role of Implicit Regularization, Batch-size and Multiple-epochs
Ayush Sekhari, Karthik Sridharan, Satyen Kale
摘要
Multi-epoch, small-batch, Stochastic Gradient Descent (SGD) has been the method of choice for learning with large over-parameterized models. A popular theory for explaining why SGD works well in practice is that the algorithm has an implicit regularization that biases its output towards a good solution. Perhaps the theoretically most well understood learning setting for SGD is that of Stochastic Convex Optimization (SCO), where it is well known that SGD learns at a rate of O(1/ p n), where n is the number of samples. In this paper, we consider the problem of SCO and explore the role of implicit regularization, batch size and multiple epochs for SGD. Our main contributions are threefold: 1. We show that for any regularizer, there is an SCO problem for which Regularized Empirical Risk Minimzation fails to learn. This automatically rules out any implicit regularization based explanation for the success of SGD. 2. We provide a separation between SGD and learning via Gradient Descent on empirical loss (GD) in terms of sample complexity. We show that there is an SCO problem such that GD with any step size and number of iterations can only learn at a suboptimal rate: at least e ⌦(1/n5/12). 3. We present a multi-epoch variant of SGD commonly used in practice. We prove that this algorithm is at least as good as single pass SGD in the worst case. However, for certain SCO problems, taking multiple passes over the dataset can significantly outperform single pass SGD. We extend our results to the general learning setting by showing a problem which is learnable for any data distribution, and for this problem, SGD is strictly better than RERM for any regularization function. We conclude by discussing the implications of our results for deep learning, and show a separation between SGD and ERM for two layer diagonal neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- On the Overlooked Structure of Stochastic GradientsZeke Xie, Qian-Yuan Tang, Mingming Sun, Ping LiNeurIPS 2023 · 被引用 18 次
- Non-convex Stochastic Composite Optimization with Polyak MomentumYuan Gao, Anton Rodomanov, Sebastian U. StichICML 2024 · 被引用 13 次
- Thinking Outside the Ball: Optimal Learning with Gradient Descent for Generalized Linear Stochastic Convex OptimizationIdan Amir, Roi Livni, Nati SrebroNeurIPS 2022 · 被引用 7 次
- From Gradient Flow on Population Loss to Learning with Stochastic Gradient DescentChristopher De Sa, Satyen Kale, Jason D. Lee, Ayush Sekhari 等NeurIPS 2022 · 被引用 5 次
- The Sample Complexity of Gradient Descent in Stochastic Convex OptimizationRoi LivniNeurIPS 2024 · 被引用 5 次
它引用的顶会 Paper2
相关 Paper
- Benign Underfitting of Stochastic Gradient DescentTomer Koren, Roi Livni, Yishay Mansour, Uri ShermanNeurIPS 2022 · 被引用 26 次
- Beyond Implicit Bias: The Insignificance of SGD Noise in Online LearningNikhil Vyas, Depen Morwani, Rosie Zhao, Gal Kaplun 等ICML 2024 · 被引用 8 次
- (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of StabilityMathieu Even, Scott Pesme, Suriya Gunasekar, Nicolas FlammarionNeurIPS 2023 · 被引用 42 次
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 被引用 235 次
- Rapid Overfitting of Multi-Pass SGD in Stochastic Convex OptimizationShira Vansover-Hager, Tomer Koren, Roi LivniICML 2025
