Lune

ICML2022顶会

Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent

Sharan Vaswani, Benjamin Dubois-Taine, Reza Babanezhad

2022年份
6顶会引用

摘要

We aim to make stochastic gradient descent (SGD) adaptive to (i) the noise σ 2 in the stochastic gradients and (ii) problemdependent constants. When minimizing smooth, strongly-convex functions with condition number κ, we prove that T iterations of SGD with exponentially decreasing step-sizes and knowledge of the smoothness can achieve an Õ exp ( -T /κ) + σ 2 /T rate, without knowing σ 2 . In order to be adaptive to the smoothness, we use a stochastic line-search (SLS) and show (via upper and lower-bounds) that SGD with SLS converges at the desired rate, but only to a neighbourhood of the solution. On the other hand, we prove that SGD with an offline estimate of the smoothness converges to the minimizer. However, its rate is slowed down proportional to the estimation error. Next, we prove that SGD with Nesterov acceleration and exponential step-sizes (referred to as ASGD) can achieve the nearoptimal Õ exp ( -T / √ κ) + σ 2 /T rate, without knowledge of σ 2 . When used with offline estimates of the smoothness and strong-convexity, ASGD still converges to the solution, albeit at a slower rate. Finally, we empirically demonstrate the effectiveness of exponential step-sizes coupled with a novel variant of SLS.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 7d9f7d84-ce2a-41e9-bcd0-29eb317a033e

引用它的顶会 Paper6

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖