Almost Bayesian: Dynamics of SGD Through Singular Learning Theory
Max Hennick, Stijn De Baerdemacker
摘要
The nature of the relationship between Bayesian sampling and stochastic gradient descent in neural networks has been a long-standing open question in the theory of deep learning. We shed light on this question by modeling the long runtime behaviour of SGD as diffusion on porous media. Using singular learning theory, we show that the late stage dynamics are strongly impacted by the degeneracies of the loss surface. From this we are able to show that under reasonable choices of hyperparameters for SGD, the local steady state distribution of SGD is effectively a tempered version of the Bayesian posterior over the weights which accounts for local accessibility constraints. We then empirically verify the porous diffusion picture across multiple models and datasets, and provide experimental evidence of the steady state-Bayesian posterior correspondence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong 等NeurIPS 2020 · 被引用 309 次
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 被引用 165 次
- Hausdorff Dimension, Heavy Tails, and Generalization in Neural NetworksUmut Simsekli, Ozan Sener, George Deligiannidis, Murat A. ErdogduNeurIPS 2020 · 被引用 79 次
- Fractal Structure and Generalization Properties of Stochastic Optimization AlgorithmsAlexander Camuto, George Deligiannidis, Murat A. Erdogdu, Mert Gürbüzbalaban 等NeurIPS 2021 · 被引用 34 次
- On Convergence-Diagnostic based Step Sizes for Stochastic Gradient DescentScott Pesme, Aymeric Dieuleveut, Nicolas FlammarionICML 2020 · 被引用 19 次
相关 Paper
- Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classificationFrancesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, Lenka ZdeborováNeurIPS 2020 · 被引用 95 次
- Three-stage Evolution and Fast Equilibrium for SGD with Non-degerate Critical PointsYi Wang, Zhiren WangICML 2022 · 被引用 4 次
- Emergence of heavy tails in homogenized stochastic gradient descentZhezhe Jiao, Martin Keller-ResselNeurIPS 2024 · 被引用 6 次
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 被引用 44 次
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen 等ICLR 2020 · 被引用 292 次
