Disentangling the Roles of Curation, Data-Augmentation and the Prior in the Cold Posterior Effect
Lorenzo Noci, Kevin Roth, Gregor Bachmann, Sebastian Nowozin, Thomas Hofmann
摘要
The "cold posterior effect" (CPE) in Bayesian deep learning describes the uncomforting observation that the predictive performance of Bayesian neural networks can be significantly improved if the Bayes posterior is artificially sharpened using a temperature parameter T < 1. The CPE is problematic in theory and practice and since the effect was identified many researchers have proposed hypotheses to explain the phenomenon. However, despite this intensive research effort the effect remains poorly understood. In this work we provide novel and nuanced evidence relevant to existing explanations for the cold posterior effect, disentangling three hypotheses: 1. The dataset curation hypothesis of Aitchison (2020): we show empirically that the CPE does not arise in a real curated data set but can be produced in a controlled experiment with varying curation strength. 2. The data augmentation hypothesis of Izmailov et al. (2021) and Fortuin et al. (2021) : we show empirically that data augmentation is sufficient but not necessary for the CPE to be present. 3. The bad prior hypothesis of Wenzel et al. ( 2020 ): we use a simple experiment evaluating the relative importance of the prior and the likelihood, strongly linking the CPE to the prior. Our results demonstrate how the CPE can arise in isolation from synthetic curation, data augmentation, and bad priors. Cold posteriors observed "in the wild" are therefore unlikely to arise from a single simple cause; as a result, we do not expect a simple "fix" for cold posteriors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- On Uncertainty, Tempering, and Data Augmentation in Bayesian ClassificationSanyam Kapoor, Wesley J. Maddox, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2022 · 被引用 64 次
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl 等ICML 2024 · 被引用 53 次
- How Tempering Fixes Data Augmentation in Bayesian Neural NetworksGregor Bachmann, Lorenzo Noci, Thomas HofmannICML 2022 · 被引用 11 次
- BayesTune: Bayesian Sparse Deep Model Fine-tuningMinyoung Kim, Timothy M. HospedalesNeurIPS 2023 · 被引用 10 次
- Monotonicity and Double Descent in Uncertainty Estimation with Gaussian ProcessesLiam Hodgkinson, Christopher van der Heide, Fred Roosta, Michael W. MahoneyICML 2023 · 被引用 9 次
它引用的顶会 Paper6
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski 等ICML 2020 · 被引用 409 次
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 被引用 362 次
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen 等ICLR 2020 · 被引用 292 次
- Bayesian Neural Network Priors RevisitedVincent Fortuin, Adrià Garriga-Alonso, Sebastian W. Ober, Florian Wenzel 等ICLR 2022 · 被引用 162 次
相关 Paper
- A statistical theory of cold posteriors in deep neural networksLaurence AitchisonICLR 2021 · 被引用 7 次
- Gaussian Mean Field Variational Inference can Overestimate Predictive VarianceJames Odgers, Ben Riegler, Siddharth Swaroop, Vincent FortuinICML 2026
- Flat Seeking Bayesian Neural NetworksVan-Anh Nguyen, Tung-Long Vuong, Hoang Phan, Thanh-Toan Do 等NeurIPS 2023 · 被引用 14 次
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 被引用 458 次
- Scalable Bayesian Learning with posteriorsSamuel Duffield, Kaelan Donatella, Johnathan Chiu, Phoebe Klett 等ICLR 2025 · 被引用 2 次
