Lune

NeurIPS2022Top-tier venue

Better SGD using Second-order Momentum

Hoang Tran, Ashok Cutkosky

2022Year
18Citations
4Top-tier citations

Abstract

We develop a new algorithm for non-convex stochastic optimization that finds an ϵ\epsilon-critical point in the optimal O(ϵ−3)O(\epsilon^{-3}) stochastic gradient and Hessian-vector product computations. Our algorithm uses Hessian-vector products to"correct"a bias term in the momentum of SGD with momentum. This leads to better gradient estimates in a manner analogous to variance reduction methods. In contrast to prior work, we do not require excessively large batch sizes, and are able to provide an adaptive algorithm whose convergence rate automatically improves with decreasing variance in the gradient estimates. We validate our results on a variety of large-scale deep learning architectures and benchmarks tasks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ca64af05-cc54-47a7-a963-eb6b9d24bfb2

Cited by top-tier papers4

Ask how each one uses it

Builds on1

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines