Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
Itamar Harel, Yonathan Wolanowsky, Gal Vardi, Nati Srebro, Daniel Soudry
Abstract
We analyze the generalization gap (gap between the training and test errors) when training a potentially over-parametrized model using a Markovian stochastic training algorithm, initialized from some distribution . We focus on Langevin dynamics with a positive temperature , i.e. gradient descent on a training loss with infinitesimal step size, perturbed with -variances Gaussian noise, and lightly regularized or bounded. There, we bound the generalization gap, at any time during training, by with probability over the dataset, where is the sample size, and with standard initialization scaling. In contrast to previous guarantees, we have no dependence on either training time or reliance on mixing, nor a dependence on dimensionality, gradient norms, or any other properties of the loss or model. This guarantee follows from a general analysis of any Markov process-based training that has a Gibbs-style stationary distribution. The proof is surprisingly simple, once we observe that the marginal distribution divergence from initialization remains bounded, as implied by a generalized second law of thermodynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29be5b2e-5c4c-43eb-93ca-08e2acc07ad7Cited by top-tier papers1
Ask how each one uses itBuilds on15
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski et al.NeurIPS 2022 · 98 citations
- On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex LearningJian Li, Xuanyuan Luo, Mingda QiaoICLR 2020 · 95 citations
- Stochastic Training is Not Necessary for GeneralizationJonas Geiping, Micah Goldblum, Phillip Pope, Michael Moeller et al.ICLR 2022 · 83 citations
Related papers
- Time-Independent Information-Theoretic Generalization Bounds for SGLDFutoshi Futami, Masahiro FujisawaNeurIPS 2023 · 12 citations
- Analyzing the Generalization Capability of SGLD Using Properties of Gaussian ChannelsHao Wang, Yizhe Huang, Rui Gao, Flávio P. CalmonNeurIPS 2021 · 32 citations
- Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvaluesJakob Kramp, Javed Lindner, Moritz HeliasICML 2026 · 2 citations
- Convergence Rates of Non-Convex Stochastic Gradient Descent Under a Generic Lojasiewicz Condition and Local SmoothnessKevin Scaman, Cédric Malherbe, Ludovic Dos SantosICML 2022 · 24 citations
- Provable Generalization of Overparameterized Meta-learning Trained with SGDYu Huang, Yingbin Liang, Longbo HuangNeurIPS 2022 · 14 citations
