Optimization and Bayes: A Trade-off for Overparameterized Neural Networks
Zhengmian Hu, Heng Huang
摘要
This paper proposes a novel algorithm, Transformative Bayesian Learning (TransBL), which bridges the gap between empirical risk minimization (ERM) and Bayesian learning for neural networks. We compare ERM, which uses gradient descent to optimize, and Bayesian learning with importance sampling for their generalization and computational complexity. We derive the first algorithm-dependent PAC-Bayesian generalization bound for infinitely wide networks based on an exact KL divergence between the trained posterior distribution obtained by infinitesimal step size gradient descent and a Gaussian prior. Moreover, we show how to transform gradient-based optimization into importance sampling by incorporating a weight. While Bayesian learning has better generalization, it suffers from low sampling efficiency. Optimization methods, on the other hand, have good sampling efficiency but poor generalization. Our proposed algorithm TransBL enables a trade-off between generalization and sampling efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 被引用 458 次
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen 等ICLR 2020 · 被引用 292 次
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 被引用 183 次
相关 Paper
- Reparameterized Importance Sampling for Robust Variational Bayesian Neural NetworksYunfei Long, Zilin Tian, Liguo Zhang, Huosheng XuICML 2024 · 被引用 1 次
- Learning under Model Misspecification: Applications to Variational and Ensemble methodsAndrés R. MasegosaNeurIPS 2020 · 被引用 112 次
- Enhancing Transfer Learning with Flexible Nonparametric Posterior SamplingHyungi Lee, Giung Nam, Edwin Fong, Juho LeeICLR 2024 · 被引用 7 次
- Scalable Bayesian Meta-Learning through Generalized Implicit GradientsYilang Zhang, Bingcong Li, Shijian Gao, Georgios B. GiannakisAAAI 2023 · 被引用 14 次
- Importance Weighted Kernel Bayes' RuleLiyuan Xu, Yutian Chen, Arnaud Doucet, Arthur GrettonICML 2022 · 被引用 5 次
