Lune

ICML2022Top-tier venue

Understanding the unstable convergence of gradient descent

Kwangjun Ahn, Jingzhao Zhang, Suvrit Sra

2022Year
89Citations
46Top-tier citations

Abstract

Most existing analyses of (stochastic) gradient descent rely on the condition that for LL-smooth costs, the step size is less than 2/L2/L. However, many works have observed that in machine learning applications step sizes often do not fulfill this condition, yet (stochastic) gradient descent still converges, albeit in an unstable manner. We investigate this unstable convergence phenomenon from first principles, and discuss key causes behind it. We also identify its main characteristics, and how they interrelate based on both theory and experiments, offering a principled view toward understanding the phenomenon.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers46

Ask how each one uses it

Builds on1

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines