What is the Long-Run Distribution of Stochastic Gradient Descent? A Large Deviations Analysis
Waïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis Mertikopoulos
摘要
In this paper, we examine the long-run distribution of stochastic gradient descent (SGD) in general, non-convex problems. Specifically, we seek to understand which regions of the problem's state space are more likely to be visited by SGD, and by how much. Using an approach based on the theory of large deviations and randomly perturbed dynamical systems, we show that the long-run distribution of SGD resembles the Boltzmann-Gibbs distribution of equilibrium thermodynamics with temperature equal to the method's step-size and energy levels determined by the problem's objective and the statistics of the noise. In particular, we show that, in the long run, (a) the problem's critical region is visited exponentially more often than any non-critical region; (b) the iterates of SGD are exponentially concentrated around the problem's minimum energy state (which does not always coincide with the global minimum of the objective); (c) all other connected components of critical points are visited with frequency that is exponentially proportional to their energy level; and, finally (d) any component of local maximizers or saddle points is "dominated" by a component of local minimizers which is visited exponentially more often.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Computing the Bias of Constant-step Stochastic Approximation with Markovian NoiseSebastian Allmeier, Nicolas GastNeurIPS 2024 · 被引用 14 次
- Temperature is All You Need for Generalization in Langevin Dynamics and other Markov ProcessesItamar Harel, Yonathan Wolanowsky, Gal Vardi, Nati Srebro 等NeurIPS 2025 · 被引用 2 次
- Multi-Agent Learning under Uncertainty: Recurrence vs. ConcentrationKyriakos Lotidis, Panayotis Mertikopoulos, Nicholas Bambos, José H. BlanchetNeurIPS 2025 · 被引用 1 次
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- Stochastic Universal Adversarial Perturbations with Fixed Optimization Constraint and Ensured High-probability TransferabilityYulin Jin, Xiaoyu Zhang, Haoyu Tong, Jian Lou 等AAAI 2026
它引用的顶会 Paper17
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 被引用 165 次
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 被引用 165 次
- What Happens after SGD Reaches Zero Loss? --A Mathematical FrameworkZhiyuan Li, Tianhao Wang, Sanjeev AroraICLR 2022 · 被引用 121 次
- On the Almost Sure Convergence of Stochastic Gradient Descent in Non-Convex ProblemsPanayotis Mertikopoulos, Nadav Hallak, Ali Kavis, Volkan CevherNeurIPS 2020 · 被引用 120 次
- The Limits of Min-Max Optimization Algorithms: Convergence to Spurious Non-Critical SetsYa-Ping Hsieh, Panayotis Mertikopoulos, Volkan CevherICML 2021 · 被引用 96 次
相关 Paper
- An Analysis of Constant Step Size SGD in the Non-convex Regime: Asymptotic Normality and BiasLu Yu, Krishnakumar Balasubramanian, Stanislav Volgushev, Murat A. ErdogduNeurIPS 2021 · 被引用 66 次
- SGD Can Converge to Local MaximaLiu Ziyin, Botao Li, James B. Simon, Masahito UedaICLR 2022 · 被引用 18 次
- Global Convergence and Stability of Stochastic Gradient DescentVivak Patel, Shushu Zhang, Bowen TianNeurIPS 2022 · 被引用 38 次
- Stochastic Gradient and Langevin ProcessesXiang Cheng, Dong Yin, Peter L. Bartlett, Michael I. JordanICML 2020 · 被引用 51 次
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala 等NeurIPS 2022 · 被引用 59 次
