Lune

ICLR2022Top-tier venue

Large Learning Rate Tames Homogeneity: Convergence and Balancing Effect

Yuqing Wang, Minshuo Chen, Tuo Zhao, Molei Tao

2022Year
53Citations
27Top-tier citations

Abstract

Recent empirical advances show that training deep models with large learning rate often improves generalization performance. However, theoretical justifications on the benefits of large learning rate are highly limited, due to challenges in analysis. In this paper, we consider using Gradient Descent (GD) with a large learning rate on a homogeneous matrix factorization problem, i.e., min⁡X,Y∥A−XY⊤∥F2\min_{X, Y} \|A - XY^\top\|_{\sf F}^2. We prove a convergence theory for constant large learning rates well beyond 2/L2/L, where LL is the largest eigenvalue of Hessian at the initialization. Moreover, we rigorously establish an implicit bias of GD induced by such a large learning rate, termed 'balancing', meaning that magnitudes of XX and YY at the limit of GD iterations will be close even if their initialization is significantly unbalanced. Numerical experiments are provided to support our theory.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers27

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines