What Are Bayesian Neural Network Posteriors Really Like?
Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon Wilson
摘要
The posterior over Bayesian neural network (BNN) parameters is extremely high-dimensional and non-convex. For computational reasons, researchers approximate this posterior using inexpensive mini-batch methods such as mean-field variational inference or stochastic-gradient Markov chain Monte Carlo (SGMCMC). To investigate foundational questions in Bayesian deep learning, we instead use full-batch Hamiltonian Monte Carlo (HMC) on modern architectures. We show that (1) BNNs can achieve significant performance gains over standard training and deep ensembles; (2) a single long HMC chain can provide a comparable representation of the posterior to multiple shorter chains; (3) in contrast to recent studies, we find posterior tempering is not needed for near-optimal performance, with little evidence for a"cold posterior"effect, which we show is largely an artifact of data augmentation; (4) BMA performance is robust to the choice of prior scale, and relatively similar for diagonal Gaussian, mixture of Gaussian, and logistic priors; (5) Bayesian neural networks show surprisingly poor generalization under domain shift; (6) while cheaper alternatives such as deep ensembles and SGMCMC methods can provide good generalization, they provide distinct predictive distributions from HMC. Notably, deep ensemble predictive distributions are similarly close to HMC as standard SGLD, and closer than standard variational inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper99
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- Uncertainty Quantification over Graph with Conformalized Graph Neural NetworksKexin Huang, Ying Jin, Emmanuel J. Candès, Jure LeskovecNeurIPS 2023 · 被引用 124 次
- Bayesian Low-rank Adaptation for Large Language ModelsAdam X. Yang, Maxime Robeyns, Xi Wang, Laurence AitchisonICLR 2024 · 被引用 111 次
它引用的顶会 Paper9
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 被引用 354 次
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen 等ICLR 2020 · 被引用 292 次
- Efficient and Scalable Bayesian Neural Nets with Rank-1 FactorsMichael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma 等ICML 2020 · 被引用 239 次
- Bayesian Neural Network Priors RevisitedVincent Fortuin, Adrià Garriga-Alonso, Sebastian W. Ober, Florian Wenzel 等ICLR 2022 · 被引用 162 次
相关 Paper
- Dangers of Bayesian Model Averaging under Covariate ShiftPavel Izmailov, Patrick Nicholson, Sanae Lotfi, Andrew Gordon WilsonNeurIPS 2021 · 被引用 51 次
- Microcanonical Langevin Ensembles: Advancing the Sampling of Bayesian Neural NetworksEmanuel Sommer, Jakob Robnik, Giorgi Nozadze, Uros Seljak 等ICLR 2025
- Posterior Refinement Improves Sample Efficiency in Bayesian Neural NetworksAgustinus Kristiadi, Runa Eschenhagen, Philipp HennigNeurIPS 2022 · 被引用 17 次
- Bayesian Posterior Approximation With Stochastic EnsemblesOleksandr Balabanov, Bernhard Mehlig, Hampus LinanderCVPR 2023
- Parameter Expanded Stochastic Gradient Markov Chain Monte CarloHyunsu Kim, Giung Nam, Chulhee Yun, Hongseok Yang 等ICLR 2025
