On the Convergence of Nesterov's Accelerated Gradient Method in Stochastic Settings
Mahmoud Assran, Mike Rabbat
Abstract
We study Nesterov's accelerated gradient method with constant step-size and momentum parameters in the stochastic approximation setting (unbiased gradients with bounded variance) and the finite-sum setting (where randomness is due to sampling mini-batches). To build better insight into the behavior of Nesterov's method in stochastic settings, we focus throughout on objectives that are smooth, strongly-convex, and twice continuously differentiable. In the stochastic approximation setting, Nesterov's method converges to a neighborhood of the optimal point at the same accelerated rate as in the deterministic setting. Perhaps surprisingly, in the finite-sum setting, we prove that Nesterov's method may diverge with the usual choice of step-size and momentum, unless additional conditions on the problem related to conditioning and data coherence are satisfied. Our results shed light as to why Nesterov's method may fail to converge or achieve acceleration in the finite-sum setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0cbd654f-fe5e-4357-910c-e1eac055d879Cited by top-tier papers12
- Stochastic Hamiltonian Gradient Methods for Smooth GamesNicolas Loizou, Hugo Berard, Alexia Jolicoeur-Martineau, Pascal Vincent et al.ICML 2020 · 54 citations
- Learning from Future: A Novel Self-Training Framework for Semantic SegmentationYe Du, Yujun Shen, Haochen Wang, Jingjing Fei et al.NeurIPS 2022 · 40 citations
- Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic ModelsCourtney Paquette, Elliot PaquetteNeurIPS 2021 · 20 citations
- Adaptive Federated Learning with Auto-Tuned ClientsJunhyung Lyle Kim, Mohammad Taha Toghani, César A. Uribe, Anastasios KyrillidisICLR 2024 · 17 citations
- An Accelerated Algorithm for Stochastic Bilevel Optimization under Unbounded SmoothnessXiaochuan Gong, Jie Hao, Mingrui LiuNeurIPS 2024 · 10 citations
Builds on1
Related papers
- Nesterov Accelerated Shuffling Gradient Method for Convex OptimizationTrang H. Tran, Katya Scheinberg, Lam M. NguyenICML 2022 · 17 citations
- Gradient correlation is a key ingredient to accelerate SGD with momentumJulien Hermant, Marien Renaud, Jean-François Aujol, Charles Dossal et al.ICLR 2025
- Nesterov acceleration in benignly non-convex landscapesKanan Gupta, Stephan WojtowytschICLR 2025
- SMG: A Shuffling Gradient-Based Method with MomentumTrang H. Tran, Lam M. Nguyen, Quoc Tran-DinhICML 2021 · 25 citations
- Unifying Nesterov's Accelerated Gradient Methods for Convex and Strongly Convex Objective FunctionsJungbin Kim, Insoon YangICML 2023 · 10 citations
