Optimization and Bayes: A Trade-off for Overparameterized Neural Networks
Zhengmian Hu, Heng Huang
Abstract
This paper proposes a novel algorithm, Transformative Bayesian Learning (TransBL), which bridges the gap between empirical risk minimization (ERM) and Bayesian learning for neural networks. We compare ERM, which uses gradient descent to optimize, and Bayesian learning with importance sampling for their generalization and computational complexity. We derive the first algorithm-dependent PAC-Bayesian generalization bound for infinitely wide networks based on an exact KL divergence between the trained posterior distribution obtained by infinitesimal step size gradient descent and a Gaussian prior. Moreover, we show how to transform gradient-based optimization into importance sampling by incorporating a weight. While Bayesian learning has better generalization, it suffers from low sampling efficiency. Optimization methods, on the other hand, have good sampling efficiency but poor generalization. Our proposed algorithm TransBL enables a trade-off between generalization and sampling efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 118939b1-503a-4f46-876a-82135bce4fa5Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen et al.ICLR 2020 · 292 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
Related papers
- Reparameterized Importance Sampling for Robust Variational Bayesian Neural NetworksYunfei Long, Zilin Tian, Liguo Zhang, Huosheng XuICML 2024 · 1 citation
- Learning under Model Misspecification: Applications to Variational and Ensemble methodsAndrés R. MasegosaNeurIPS 2020 · 112 citations
- Enhancing Transfer Learning with Flexible Nonparametric Posterior SamplingHyungi Lee, Giung Nam, Edwin Fong, Juho LeeICLR 2024 · 7 citations
- Scalable Bayesian Meta-Learning through Generalized Implicit GradientsYilang Zhang, Bingcong Li, Shijian Gao, Georgios B. GiannakisAAAI 2023 · 14 citations
- Importance Weighted Kernel Bayes' RuleLiyuan Xu, Yutian Chen, Arnaud Doucet, Arthur GrettonICML 2022 · 5 citations
