Beyond Deep Ensembles: A Large-Scale Evaluation of Bayesian Deep Learning under Distribution Shift
Florian Seligmann, Philipp Becker, Michael Volpp, Gerhard Neumann
Abstract
Bayesian deep learning (BDL) is a promising approach to achieve well-calibrated predictions on distribution-shifted data. Nevertheless, there exists no large-scale survey that evaluates recent SOTA methods on diverse, realistic, and challenging benchmark tasks in a systematic manner. To provide a clear picture of the current state of BDL research, we evaluate modern BDL algorithms on real-world datasets from the WILDS collection containing challenging classification and regression tasks, with a focus on generalization capability and calibration under distribution shift. We compare the algorithms on a wide range of large, convolutional and transformer-based neural network architectures. In particular, we investigate a signed version of the expected calibration error that reveals whether the methods are over- or under-confident, providing further insight into the behavior of the methods. Further, we provide the first systematic evaluation of BDL for fine-tuning large pre-trained models, where training from scratch is prohibitively expensive. Finally, given the recent success of Deep Ensembles, we extend popular single-mode posterior approximations to multiple modes by the use of ensembles. While we find that ensembling single-mode approximations generally improves the generalization capability and calibration of the models by a significant margin, we also identify a failure mode of ensembles when finetuning large transformer-based language models. In this setting, variational inference based approaches such as last-layer Bayes By Backprop outperform other methods in terms of accuracy by a large margin, while modern approximate inference algorithms such as SWAG achieve the best calibration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f033c46-0627-41f0-96d1-672ebe56e9ecCited by top-tier papers5
- BTFL: A Bayesian-based Test-Time Generalization Method for Internal and External Data Distributions in Federated learningYu Zhou, Bingyan LiuKDD 2025 · 3 citations
- Bayesian Low-Rank Learning (Bella): A Practical Approach to Bayesian Neural NetworksBao Gia Doan, Afshar Shamsi, Xiao-Yu Guo, Arash Mohammadi et al.AAAI 2025 · 1 citation
- Target-Agnostic Calibration under Distribution Shift with Frequency-Aware Gradient RectificationYilin Zhang, Cai Xu, You Wu, Ziyu Guan et al.ICML 2026 · 1 citation
- The Disparate Benefits of Deep EnsemblesKajetan Schweighofer, Adrián Arnaiz-Rodríguez, Sepp Hochreiter, Nuria OliverICML 2025
- Stop Diverse OOD Attacks: Knowledge Ensemble for Reliable DefenseZhenbo Shi, Xiaoman Liu, Yuxuan Zhang, Shuchang Wang et al.AAAI 2025
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran et al.NeurIPS 2020 · 604 citations
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
Related papers
- Exploring the Uncertainty Properties of Neural Networks' Implicit Priors in the Infinite-Width LimitBen Adlam, Jaehoon Lee, Lechao Xiao, Jeffrey Pennington et al.ICLR 2021 · 3 citations
- Variational Bayesian Last LayersJames Harrison, John Willes, Jasper SnoekICLR 2024 · 75 citations
- Do Bayesian Neural Networks Actually Behave Like Bayesian Models?Gábor Pituk, Vik Shirvaikar, Tom RainforthICML 2025
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 344 citations
- On the Expressiveness of Approximate Inference in Bayesian Neural NetworksAndrew Y. K. Foong, David R. Burt, Yingzhen Li, Richard E. TurnerNeurIPS 2020 · 142 citations
