Robustness Analysis of Non-Convex Stochastic Gradient Descent using Biased Expectations
Kevin Scaman, Cédric Malherbe
摘要
This work proposes a novel analysis of stochastic gradient descent (SGD) for non-convex and smooth optimization. Our analysis sheds light on the impact of the probability distribution of the gradient noise on the convergence rate of the norm of the gradient. In the case of sub-Gaussian and centered noise, we prove that, with probability 1δ, the number of iterations to reach a precision ε for the squared gradient norm is O(ε -2 ln(1/δ)). In the case of centered and integrable heavytailed noise, we show that, while the expectation of the iterates may be infinite, the squared gradient norm still converges with probability 1δ in O(ε -p δ -q ) iterations, where p, q > 2. This result shows that heavy-tailed noise on the gradient slows down the convergence of SGD without preventing it, proving that SGD is robust to gradient noise with unbounded variance, a setting of interest for Deep Learning. In addition, it indicates that choosing a step size proportional to T -1/b where b is the tail-parameter of the noise and T is the number of iterations leads to the best convergence rates. Both results are simple corollaries of a unified analysis using the novel concept of biased expectations, a simple and intuitive mathematical tool to obtain concentration inequalities. Using this concept, we propose a new quantity to measure the amount of noise added to the gradient, and discuss its value in multiple scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- High-probability Bounds for Non-Convex Stochastic Optimization with Heavy TailsAshok Cutkosky, Harsh MehtaNeurIPS 2021 · 被引用 119 次
- Convergence Rates of Non-Convex Stochastic Gradient Descent Under a Generic Lojasiewicz Condition and Local SmoothnessKevin Scaman, Cédric Malherbe, Ludovic Dos SantosICML 2022 · 被引用 24 次
- Globally Convergent Policy Search for Output EstimationJack Umenberger, Max Simchowitz, Juan C. Perdomo, Kaiqing Zhang 等NeurIPS 2022 · 被引用 16 次
- Cautious Weight DecayLizhang Chen, Jonathan Li, Kaizhao Liang, Baiyu Su 等ICLR 2026 · 被引用 14 次
- Existence and Estimation of Critical Batch Size for Training Generative Adversarial Networks with Two Time-Scale Update RuleNaoki Sato, Hideaki IidukaICML 2023 · 被引用 11 次
它引用的顶会 Paper1
相关 Paper
- Stochastic Gradient Methods under Heavy-Tailed Noises in Weakly Convex OptimizationTianxi Zhu, Yi Xu, Qi Wang, Xiangyang JiICML 2026
- High Probability Guarantees for Nonconvex Stochastic Gradient Descent with Heavy TailsShaojie Li, Yong LiuICML 2022 · 被引用 37 次
- Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined AnalysisZijian LiuICLR 2026 · 被引用 5 次
- Revisiting the Last-Iterate Convergence of Stochastic Gradient MethodsZijian Liu, Zhengyuan ZhouICLR 2024 · 被引用 32 次
- Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-TailedSavelii Chezhegov, Yaroslav Klyukin, Andrei Semenov, Aleksandr Beznosikov 等ICML 2025
