Disentangling the Roles of Curation, Data-Augmentation and the Prior in the Cold Posterior Effect
Lorenzo Noci, Kevin Roth, Gregor Bachmann, Sebastian Nowozin, Thomas Hofmann
Abstract
The "cold posterior effect" (CPE) in Bayesian deep learning describes the uncomforting observation that the predictive performance of Bayesian neural networks can be significantly improved if the Bayes posterior is artificially sharpened using a temperature parameter T < 1. The CPE is problematic in theory and practice and since the effect was identified many researchers have proposed hypotheses to explain the phenomenon. However, despite this intensive research effort the effect remains poorly understood. In this work we provide novel and nuanced evidence relevant to existing explanations for the cold posterior effect, disentangling three hypotheses: 1. The dataset curation hypothesis of Aitchison (2020): we show empirically that the CPE does not arise in a real curated data set but can be produced in a controlled experiment with varying curation strength. 2. The data augmentation hypothesis of Izmailov et al. (2021) and Fortuin et al. (2021) : we show empirically that data augmentation is sufficient but not necessary for the CPE to be present. 3. The bad prior hypothesis of Wenzel et al. ( 2020 ): we use a simple experiment evaluating the relative importance of the prior and the likelihood, strongly linking the CPE to the prior. Our results demonstrate how the CPE can arise in isolation from synthetic curation, data augmentation, and bad priors. Cold posteriors observed "in the wild" are therefore unlikely to arise from a single simple cause; as a result, we do not expect a simple "fix" for cold posteriors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4319d28f-d8a7-4a0e-9777-39dd54cefe9dCited by top-tier papers10
- On Uncertainty, Tempering, and Data Augmentation in Bayesian ClassificationSanyam Kapoor, Wesley J. Maddox, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2022 · 64 citations
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl et al.ICML 2024 · 53 citations
- How Tempering Fixes Data Augmentation in Bayesian Neural NetworksGregor Bachmann, Lorenzo Noci, Thomas HofmannICML 2022 · 11 citations
- BayesTune: Bayesian Sparse Deep Model Fine-tuningMinyoung Kim, Timothy M. HospedalesNeurIPS 2023 · 10 citations
- Monotonicity and Double Descent in Uncertainty Estimation with Gaussian ProcessesLiam Hodgkinson, Christopher van der Heide, Fred Roosta, Michael W. MahoneyICML 2023 · 9 citations
Builds on6
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen et al.ICLR 2020 · 292 citations
- Bayesian Neural Network Priors RevisitedVincent Fortuin, Adrià Garriga-Alonso, Sebastian W. Ober, Florian Wenzel et al.ICLR 2022 · 162 citations
Related papers
- A statistical theory of cold posteriors in deep neural networksLaurence AitchisonICLR 2021 · 7 citations
- Gaussian Mean Field Variational Inference can Overestimate Predictive VarianceJames Odgers, Ben Riegler, Siddharth Swaroop, Vincent FortuinICML 2026
- Flat Seeking Bayesian Neural NetworksVan-Anh Nguyen, Tung-Long Vuong, Hoang Phan, Thanh-Toan Do et al.NeurIPS 2023 · 14 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- Scalable Bayesian Learning with posteriorsSamuel Duffield, Kaelan Donatella, Johnathan Chiu, Phoebe Klett et al.ICLR 2025 · 2 citations
