Power-Law Escape Rate of SGD
Takashi Mori, Liu Ziyin, Kangqiao Liu, Masahito Ueda
Abstract
Stochastic gradient descent (SGD) undergoes complicated multiplicative noise for the mean-square loss. We use this property of SGD noise to derive a stochastic differential equation (SDE) with simpler additive noise by performing a random time change. Using this formalism, we show that the log loss barrier between a local minimum and a saddle determines the escape rate of SGD from the local minimum, contrary to the previous results borrowing from physics that the linear loss barrier decides the escape rate. Our escape-rate formula strongly depends on the typical magnitude and the number of the outlier eigenvalues of the Hessian. This result explains an empirical fact that SGD prefers flat minima with low effective dimensions, giving an insight into implicit biases of SGD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf4d7540-b4f0-4188-be23-69b4372fc960Cited by top-tier papers17
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 101 citations
- The Implicit Regularization of Dynamical Stability in Stochastic Gradient DescentLei Wu, Weijie J. SuICML 2023 · 41 citations
- Exact Solutions of a Deep Linear NetworkLiu Ziyin, Botao Li, Xiangming MengNeurIPS 2022 · 31 citations
- What is the Long-Run Distribution of Stochastic Gradient Descent? A Large Deviations AnalysisWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2024 · 17 citations
- Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling LawsJinbo Wang, Binghui Li, Zhanpeng Zhou, Mingze Wang et al.ICLR 2026 · 6 citations
Builds on4
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 165 citations
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 165 citations
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan et al.ICML 2020 · 125 citations
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient DescentKangqiao Liu, Liu Ziyin, Masahito UedaICML 2021 · 46 citations
Related papers
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- The alignment property of SGD noise and how it helps select flat minima: A stability analysisLei Wu, Mingze Wang, Weijie SuNeurIPS 2022 · 80 citations
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong et al.NeurIPS 2020 · 309 citations
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 44 citations
- What Happens after SGD Reaches Zero Loss? --A Mathematical FrameworkZhiyuan Li, Tianhao Wang, Sanjeev AroraICLR 2022 · 121 citations
