Lune

NeurIPS2023顶会

Two Sides of One Coin: the Limits of Untuned SGD and the Power of Adaptive Methods

Junchi Yang, Xiang Li, Ilyas Fatkhullin, Niao He

2023年份
33被引次数
9顶会引用

摘要

The classical analysis of Stochastic Gradient Descent (SGD) with polynomially decaying stepsize ηt=η/t\eta_t = \eta/\sqrt{t} relies on well-tuned η\eta depending on problem parameters such as Lipschitz smoothness constant, which is often unknown in practice. In this work, we prove that SGD with arbitrary η>0\eta>0, referred to as untuned SGD, still attains an order-optimal convergence rate O~(T−1/4)\widetilde{O}(T^{-1/4}) in terms of gradient norm for minimizing smooth objectives. Unfortunately, it comes at the expense of a catastrophic exponential dependence on the smoothness constant, which we show is unavoidable for this scheme even in the noiseless setting. We then examine three families of adaptive methods \unicodex2013\unicode{x2013} Normalized SGD (NSGD), AMSGrad, and AdaGrad \unicodex2013\unicode{x2013} unveiling their power in preventing such exponential dependency in the absence of information about the smoothness parameter and boundedness of stochastic gradients. Our results provide theoretical justification for the advantage of adaptive methods over untuned SGD in alleviating the issue with large gradients.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper9

问问它们各自怎么用它

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖