Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
Itamar Harel, Yonathan Wolanowsky, Gal Vardi, Nati Srebro, Daniel Soudry
摘要
We analyze the generalization gap (gap between the training and test errors) when training a potentially over-parametrized model using a Markovian stochastic training algorithm, initialized from some distribution . We focus on Langevin dynamics with a positive temperature , i.e. gradient descent on a training loss with infinitesimal step size, perturbed with -variances Gaussian noise, and lightly regularized or bounded. There, we bound the generalization gap, at any time during training, by with probability over the dataset, where is the sample size, and with standard initialization scaling. In contrast to previous guarantees, we have no dependence on either training time or reliance on mixing, nor a dependence on dimensionality, gradient norms, or any other properties of the loss or model. This guarantee follows from a general analysis of any Markov process-based training that has a Gibbs-style stationary distribution. The proof is surprisingly simple, once we observe that the marginal distribution divergence from initialization remains bounded, as implied by a generalized second law of thermodynamics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski 等ICML 2020 · 被引用 409 次
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski 等NeurIPS 2022 · 被引用 98 次
- On Generalization Error Bounds of Noisy Gradient Methods for Non-Convex LearningJian Li, Xuanyuan Luo, Mingda QiaoICLR 2020 · 被引用 95 次
- Stochastic Training is Not Necessary for GeneralizationJonas Geiping, Micah Goldblum, Phillip Pope, Michael Moeller 等ICLR 2022 · 被引用 83 次
相关 Paper
- Time-Independent Information-Theoretic Generalization Bounds for SGLDFutoshi Futami, Masahiro FujisawaNeurIPS 2023 · 被引用 12 次
- Analyzing the Generalization Capability of SGLD Using Properties of Gaussian ChannelsHao Wang, Yizhe Huang, Rui Gao, Flávio P. CalmonNeurIPS 2021 · 被引用 32 次
- Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvaluesJakob Kramp, Javed Lindner, Moritz HeliasICML 2026 · 被引用 2 次
- Convergence Rates of Non-Convex Stochastic Gradient Descent Under a Generic Lojasiewicz Condition and Local SmoothnessKevin Scaman, Cédric Malherbe, Ludovic Dos SantosICML 2022 · 被引用 24 次
- Provable Generalization of Overparameterized Meta-learning Trained with SGDYu Huang, Yingbin Liang, Longbo HuangNeurIPS 2022 · 被引用 14 次
