Lune

ICLR2021Top-tier venue

RMSprop converges with proper hyper-parameter

Naichen Shi, Dawei Li, Mingyi Hong, Ruoyu Sun

2021Year
79Citations
24Top-tier citations

Abstract

Despite the existence of divergence examples, RMSprop remains one of the most popular algorithms in machine learning. Towards closing the gap between theory and practice, we prove that RMSprop converges with proper choice of hyper-parameters under certain conditions. More specifically, we prove that when the hyper-parameter β2\beta_2 is close enough to 11, RMSprop and its random shuffling version converge to a bounded region in general, and to critical points in the interpolation regime. It is worth mentioning that our results do not depend on ``bounded gradient" assumption, which is often the key assumption utilized by existing theoretical work for Adam-type adaptive gradient method. Removing this assumption allows us to establish a phase transition from divergence to non-divergence for RMSprop. Finally, based on our theory, we conjecture that in practice there is a critical threshold β2∗\sf{\beta_2^*}, such that RMSprop generates reasonably good results only if 1>β2≥β2∗1>\beta_2\ge \sf{\beta_2^*}. We provide empirical evidence for such a phase transition in our numerical experiments.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get da597250-bdda-4457-a22b-b745a356a396

Cited by top-tier papers24

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines