Uncertainty Quantification and Deep Ensembles
Rahul Rahaman, Alexandre H. Thiéry
Abstract
Deep Learning methods are known to suffer from calibration issues: they typically produce over-confident estimates. These problems are exacerbated in the low data regime. Although the calibration of probabilistic models is well studied, calibrating extremely over-parametrized models in the low-data regime presents unique challenges. We show that deep-ensembles do not necessarily lead to improved calibration properties. In fact, we show that standard ensembling methods, when used in conjunction with modern techniques such as mixup regularization, can lead to less calibrated models. In this text, we examine the interplay between three of the most simple and commonly used approaches to leverage deep learning when data is scarce: data-augmentation, ensembling, and post-processing calibration methods. We demonstrate that, although standard ensembling techniques certainly help to boost accuracy, the calibration of deep-ensembles relies on subtle trade-offs. Our main finding is that calibration methods such as temperature scaling need to be slightly tweaked when used with deep-ensembles and, crucially, need to be executed after the averaging process. Our simulations indicate that, in the low data regime, this simple strategy can halve the Expected Calibration Error (ECE) on a range of benchmark classification problems when compared to standard deep-ensembles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed49edff-831f-40d9-875b-13a678d7c7c3Cited by top-tier papers29
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
- Deep Ensembles Work, But Are They Necessary?Taiga Abe, Estefany Kelly Buchanan, Geoff Pleiss, Richard S. Zemel et al.NeurIPS 2022 · 101 citations
- GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual ReasoningJusheng Zhang, Yijia Fan, Wenjun Lin, Ruiqi Chen et al.NeurIPS 2025 · 75 citations
- Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution RobustnessFrancesco Pinto, Harry Yang, Ser Nam Lim, Philip H. S. Torr et al.NeurIPS 2022 · 74 citations
- Combining Ensembles and Data Augmentation Can Harm Your CalibrationYeming Wen, Ghassen Jerfel, Rafael Muller, Michael W. Dusenberry et al.ICLR 2021 · 72 citations
Builds on4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- Combining Ensembles and Data Augmentation Can Harm Your CalibrationYeming Wen, Ghassen Jerfel, Rafael Muller, Michael W. Dusenberry et al.ICLR 2021 · 72 citations
Related papers
- On the Limitations of Temperature Scaling for Distributions with OverlapsMuthu Chidambaram, Rong GeICLR 2024 · 11 citations
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 177 citations
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 276 citations
- On the Pitfall of Mixup for Uncertainty CalibrationDeng-Bao Wang, Lanqing Li, Peilin Zhao, Pheng-Ann Heng et al.CVPR 2023
- Ex Uno Pluria: Insights on Ensembling in Low Precision Number SystemsGiung Nam, Juho LeeNeurIPS 2024 · 2 citations
