Accelerating SGD with momentum for over-parameterized learning
Chaoyue Liu, Mikhail Belkin
摘要
Nesterov SGD is widely used for training modern neural networks and other machine learning models. Yet, its advantages over SGD have not been theoretically clarified. Indeed, as we show in our paper, both theoretically and empirically, Nesterov SGD with any parameter selection does not in general provide acceleration over ordinary SGD. Furthermore, Nesterov SGD may diverge for step sizes that ensure convergence of ordinary SGD. This is in contrast to the classical results in the deterministic scenario, where the same step size ensures accelerated convergence of the Nesterov's method over optimal gradient descent. To address the non-acceleration issue, we introduce a compensation term to Nesterov SGD. The resulting algorithm, which we call MaSS, converges for same step sizes as SGD. We prove that MaSS obtains an accelerated convergence rates over SGD for any mini-batch size in the linear setting. For full batch, the convergence rate of MaSS matches the well-known accelerated rate of the Nesterov's method. We also analyze the practically important question of the dependence of the convergence rate and optimal hyper-parameters on the mini-batch size, demonstrating three distinct regimes: linear scaling, diminishing returns and saturation. Experimental evaluation of MaSS for several standard architectures of deep networks, including ResNet and convolutional networks, shows improved performance over SGD, Nesterov SGD and Adam.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
- Training Neural Networks for and by InterpolationLeonard Berrada, Andrew Zisserman, M. Pawan KumarICML 2020 · 被引用 71 次
- On the Convergence of Nesterov's Accelerated Gradient Method in Stochastic SettingsMahmoud Assran, Mike RabbatICML 2020 · 被引用 71 次
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 52 次
- First Order Methods with Markovian Noise: from Acceleration to Variational InequalitiesAleksandr Beznosikov, Sergey Samsonov, Marina Sheshukova, Alexander V. Gasnikov 等NeurIPS 2023 · 被引用 26 次
相关 Paper
- Win: Weight-Decay-Integrated Nesterov Acceleration for Adaptive Gradient AlgorithmsPan Zhou, Xingyu Xie, Shuicheng YanICLR 2023
- Gradient correlation is a key ingredient to accelerate SGD with momentumJulien Hermant, Marien Renaud, Jean-François Aujol, Charles Dossal 等ICLR 2025
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 被引用 25 次
- Momentum-SAM: Sharpness Aware Minimization without Computational OverheadMarlon Becker, Frederick Altrock, Benjamin RisseNeurIPS 2025 · 被引用 16 次
- Large Batch Optimization for Deep Learning Using New Complete Layer-Wise Adaptive Rate ScalingZhouyuan Huo, Bin Gu, Heng HuangAAAI 2021 · 被引用 35 次
