Deep Ensembles Work, But Are They Necessary?
Taiga Abe, Estefany Kelly Buchanan, Geoff Pleiss, Richard S. Zemel, John P. Cunningham
Abstract
Ensembling neural networks is an effective way to increase accuracy, and can often match the performance of individual larger models. This observation poses a natural question: given the choice between a deep ensemble and a single neural network with similar accuracy, is one preferable over the other? Recent work suggests that deep ensembles may offer distinct benefits beyond predictive power: namely, uncertainty quantification and robustness to dataset shift. In this work, we demonstrate limitations to these purported benefits, and show that a single (but larger) neural network can replicate these qualities. First, we show that ensemble diversity, by any metric, does not meaningfully contribute to an ensemble's uncertainty quantification on out-of-distribution (OOD) data, but is instead highly correlated with the relative improvement of a single larger model. Second, we show that the OOD performance afforded by ensembles is strongly determined by their in-distribution (InD) performance, and-in this sense-is not indicative of any "effective robustness." While deep ensembles are a practical way to achieve improvements to predictive power, uncertainty quantification, and robustness, our results show that these improvements can be replicated by a (larger) single model. Recent research suggests that deep ensembles may be preferable to single models in safety-critical applications and settings where data shifts significantly away from the training distribution. First, Lakshminarayanan et al. [45] demonstrate that deep ensembles provide well-calibrated estimates of * Equal contribution. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a746ccb0-19f6-452c-8367-9395eeb89e6eCited by top-tier papers13
- When are ensembles really effective?Ryan Theisen, Hyunsuk Kim, Yaoqing Yang, Liam Hodgkinson et al.NeurIPS 2023 · 29 citations
- Out-of-Distribution Detection via Deep Multi-Comprehension EnsembleChenhui Xu, Fuxun Yu, Zirui Xu, Nathan Inkawhich et al.ICML 2024 · 14 citations
- FedCal: Achieving Local and Global Calibration in Federated Learning via Aggregated Parameterized ScalerHongyi Peng, Han Yu, Xiaoli Tang, Xiaoxiao LiICML 2024 · 12 citations
- Uncertainty Quantification with the Empirical Neural Tangent KernelJoseph Wilson, Chris van der Heide, Liam Hodgkinson, Fred RoostaNeurIPS 2025 · 11 citations
- Relaxed Quantile Regression: Prediction Intervals for Asymmetric NoiseThomas Pouplin, Alan Jeffares, Nabeel Seedat, Mihaela van der SchaarICML 2024 · 9 citations
Builds on18
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran et al.NeurIPS 2020 · 604 citations
Related papers
- Neural Ensemble Search for Uncertainty Estimation and Dataset ShiftSheheryar Zaidi, Arber Zela, Thomas Elsken, Chris C. Holmes et al.NeurIPS 2021 · 97 citations
- Theoretical Limitations of Ensembles in the Age of OverparameterizationNiclas Dern, John Patrick Cunningham, Geoff PleissICML 2025
- Ensemble Distribution DistillationAndrey Malinin, Bruno Mlodozeniec, Mark J. F. GalesICLR 2020 · 273 citations
- Robustness via Cross-Domain EnsemblesTeresa Yeo, Oguzhan Fatih Kar, Amir ZamirICCV 2021 · 30 citations
- On Power Laws in Deep EnsemblesEkaterina Lobacheva, Nadezhda Chirkova, Maxim Kodryan, Dmitry P. VetrovNeurIPS 2020 · 48 citations
