Towards understanding how momentum improves generalization in deep learning
Samy Jelassi, Yuanzhi Li
Abstract
Stochastic gradient descent (SGD) with momentum is widely used for training modern deep learning architectures. While it is well-understood that using momentum can lead to faster convergence rate in various settings, it has also been observed that momentum yields higher generalization. Prior work argue that momentum stabilizes the SGD noise during training and this leads to higher generalization. In this paper, we adopt another perspective and first empirically show that gradient descent with momentum (GD+M) significantly improves generalization compared to gradient descent (GD) in some deep learning problems. From this observation, we formally study how momentum improves generalization. We devise a binary classification setting where a one-hidden layer (over-parameterized) convolutional neural network trained with GD+M provably generalizes better than the same network trained with GD, when both algorithms are similarly initialized. The key insight in our analysis is that momentum is beneficial in datasets where the examples share some feature but differ in their margin. Contrary to GD that memorizes the small margin data, GD+M still learns the feature in these data thanks to its historical gradients. Lastly, we empirically validate our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9941380e-b572-4b8f-90eb-c2b2d224bc7dCited by top-tier papers35
- Vision Transformers provably learn spatial structureSamy Jelassi, Michael E. Sander, Yuanzhi LiNeurIPS 2022 · 115 citations
- Flora: Low-Rank Adapters Are Secretly Gradient CompressorsYongchang Hao, Yanshuai Cao, Lili MouICML 2024 · 113 citations
- Robust Learning with Progressive Data Expansion Against Spurious CorrelationYihe Deng, Yu Yang, Baharan Mirzasoleiman, Quanquan GuNeurIPS 2023 · 53 citations
- Momentum Provably Improves Error Feedback!Ilyas Fatkhullin, Alexander Tyurin, Peter RichtárikNeurIPS 2023 · 47 citations
- The Mechanism of Prediction Head in Non-contrastive Self-supervised LearningZixin Wen, Yuanzhi LiNeurIPS 2022 · 44 citations
Builds on6
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- An Improved Analysis of Stochastic Gradient Descent with MomentumYanli Liu, Yuan Gao, Wotao YinNeurIPS 2020 · 328 citations
- Momentum Improves Normalized SGDAshok Cutkosky, Harsh MehtaICML 2020 · 177 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- The Implicit and Explicit Regularization Effects of DropoutColin Wei, Sham M. Kakade, Tengyu MaICML 2020 · 129 citations
Related papers
- The Marginal Value of Momentum for Small Learning Rate SGDRunzhe Wang, Sadhika Malladi, Tianhao Wang, Kaifeng Lyu et al.ICLR 2024 · 14 citations
- Demystify Hyperparameters for Stochastic Optimization with Transferable RepresentationsJianhui Sun, Mengdi Huai, Kishlay Jha, Aidong ZhangKDD 2022 · 5 citations
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 25 citations
- Does Momentum Change the Implicit Regularization on Separable Data?Bohan Wang, Qi Meng, Huishuai Zhang, Ruoyu Sun et al.NeurIPS 2022 · 29 citations
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear NetworkJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICML 2021 · 26 citations
