Lune

NeurIPS2023Top-tier venue

Two Sides of One Coin: the Limits of Untuned SGD and the Power of Adaptive Methods

Junchi Yang, Xiang Li, Ilyas Fatkhullin, Niao He

2023Year
33Citations
9Top-tier citations

Abstract

The classical analysis of Stochastic Gradient Descent (SGD) with polynomially decaying stepsize ηt=η/t\eta_t = \eta/\sqrt{t} relies on well-tuned η\eta depending on problem parameters such as Lipschitz smoothness constant, which is often unknown in practice. In this work, we prove that SGD with arbitrary η>0\eta>0, referred to as untuned SGD, still attains an order-optimal convergence rate O~(T−1/4)\widetilde{O}(T^{-1/4}) in terms of gradient norm for minimizing smooth objectives. Unfortunately, it comes at the expense of a catastrophic exponential dependence on the smoothness constant, which we show is unavoidable for this scheme even in the noiseless setting. We then examine three families of adaptive methods \unicodex2013\unicode{x2013} Normalized SGD (NSGD), AMSGrad, and AdaGrad \unicodex2013\unicode{x2013} unveiling their power in preventing such exponential dependency in the absence of information about the smoothness parameter and boundedness of stochastic gradients. Our results provide theoretical justification for the advantage of adaptive methods over untuned SGD in alleviating the issue with large gradients.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext dc615cd0-77f8-49d4-b3cf-a0101304bba4

Cited by top-tier papers9

Ask how each one uses it

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines