What is the Long-Run Distribution of Stochastic Gradient Descent? A Large Deviations Analysis
Waïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis Mertikopoulos
Abstract
In this paper, we examine the long-run distribution of stochastic gradient descent (SGD) in general, non-convex problems. Specifically, we seek to understand which regions of the problem's state space are more likely to be visited by SGD, and by how much. Using an approach based on the theory of large deviations and randomly perturbed dynamical systems, we show that the long-run distribution of SGD resembles the Boltzmann-Gibbs distribution of equilibrium thermodynamics with temperature equal to the method's step-size and energy levels determined by the problem's objective and the statistics of the noise. In particular, we show that, in the long run, (a) the problem's critical region is visited exponentially more often than any non-critical region; (b) the iterates of SGD are exponentially concentrated around the problem's minimum energy state (which does not always coincide with the global minimum of the objective); (c) all other connected components of critical points are visited with frequency that is exponentially proportional to their energy level; and, finally (d) any component of local maximizers or saddle points is "dominated" by a component of local minimizers which is visited exponentially more often.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbbfeefe-ab34-4d8b-a6f4-1993cd0a3eb6Cited by top-tier papers6
- Computing the Bias of Constant-step Stochastic Approximation with Markovian NoiseSebastian Allmeier, Nicolas GastNeurIPS 2024 · 14 citations
- Temperature is All You Need for Generalization in Langevin Dynamics and other Markov ProcessesItamar Harel, Yonathan Wolanowsky, Gal Vardi, Nati Srebro et al.NeurIPS 2025 · 2 citations
- Multi-Agent Learning under Uncertainty: Recurrence vs. ConcentrationKyriakos Lotidis, Panayotis Mertikopoulos, Nicholas Bambos, José H. BlanchetNeurIPS 2025 · 1 citation
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- Stochastic Universal Adversarial Perturbations with Fixed Optimization Constraint and Ensured High-probability TransferabilityYulin Jin, Xiaoyu Zhang, Haoyu Tong, Jian Lou et al.AAAI 2026
Builds on17
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 165 citations
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 165 citations
- What Happens after SGD Reaches Zero Loss? --A Mathematical FrameworkZhiyuan Li, Tianhao Wang, Sanjeev AroraICLR 2022 · 121 citations
- On the Almost Sure Convergence of Stochastic Gradient Descent in Non-Convex ProblemsPanayotis Mertikopoulos, Nadav Hallak, Ali Kavis, Volkan CevherNeurIPS 2020 · 120 citations
- The Limits of Min-Max Optimization Algorithms: Convergence to Spurious Non-Critical SetsYa-Ping Hsieh, Panayotis Mertikopoulos, Volkan CevherICML 2021 · 96 citations
Related papers
- An Analysis of Constant Step Size SGD in the Non-convex Regime: Asymptotic Normality and BiasLu Yu, Krishnakumar Balasubramanian, Stanislav Volgushev, Murat A. ErdogduNeurIPS 2021 · 66 citations
- SGD Can Converge to Local MaximaLiu Ziyin, Botao Li, James B. Simon, Masahito UedaICLR 2022 · 18 citations
- Global Convergence and Stability of Stochastic Gradient DescentVivak Patel, Shushu Zhang, Bowen TianNeurIPS 2022 · 38 citations
- Stochastic Gradient and Langevin ProcessesXiang Cheng, Dong Yin, Peter L. Bartlett, Michael I. JordanICML 2020 · 51 citations
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala et al.NeurIPS 2022 · 59 citations
