Large Learning Rate Tames Homogeneity: Convergence and Balancing Effect
Yuqing Wang, Minshuo Chen, Tuo Zhao, Molei Tao
摘要
Recent empirical advances show that training deep models with large learning rate often improves generalization performance. However, theoretical justifications on the benefits of large learning rate are highly limited, due to challenges in analysis. In this paper, we consider using Gradient Descent (GD) with a large learning rate on a homogeneous matrix factorization problem, i.e., . We prove a convergence theory for constant large learning rates well beyond , where is the largest eigenvalue of Hessian at the initialization. Moreover, we rigorously establish an implicit bias of GD induced by such a large learning rate, termed 'balancing', meaning that magnitudes of and at the limit of GD iterations will be close even if their initialization is significantly unbalanced. Numerical experiments are provided to support our theory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 被引用 139 次
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- SGD with Large Step Sizes Learns Sparse FeaturesMaksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionICML 2023 · 被引用 77 次
- Implicit Bias of Gradient Descent for Logistic Regression at the Edge of StabilityJingfeng Wu, Vladimir Braverman, Jason D. LeeNeurIPS 2023 · 被引用 46 次
- (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of StabilityMathieu Even, Scott Pesme, Suriya Gunasekar, Nicolas FlammarionNeurIPS 2023 · 被引用 42 次
它引用的顶会 Paper6
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong 等NeurIPS 2020 · 被引用 309 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- Global Convergence of Gradient Descent for Asymmetric Low-Rank Matrix FactorizationTian Ye, Simon S. DuNeurIPS 2021 · 被引用 61 次
- Implicit Bias of Gradient Descent based Adversarial Training on Separable DataYan Li, Ethan X. Fang, Huan Xu, Tuo ZhaoICLR 2020 · 被引用 40 次
相关 Paper
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 被引用 155 次
- Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence RegionShuang Liang, Guido MontufarICLR 2026 · 被引用 5 次
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 被引用 53 次
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 被引用 178 次
- Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional DataSpencer Frei, Gal Vardi, Peter L. Bartlett, Nathan Srebro 等ICLR 2023 · 被引用 5 次
