Better Uncertainty Calibration via Proper Scores for Classification and Beyond
Sebastian G. Gruber, Florian Buettner
Abstract
With model trustworthiness being crucial for sensitive real-world applications, practitioners are putting more and more focus on improving the uncertainty calibration of deep neural networks. Calibration errors are designed to quantify the reliability of probabilistic predictions but their estimators are usually biased and inconsistent. In this work, we introduce the framework of proper calibration errors, which relates every calibration error to a proper score and provides a respective upper bound with optimal estimation properties. This relationship can be used to reliably quantify the model calibration improvement. We theoretically and empirically demonstrate the shortcomings of commonly used estimators compared to our approach. Due to the wide applicability of proper scores, this gives a natural extension of recalibration beyond classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Dual Focal Loss for CalibrationLinwei Tao, Minjing Dong, Chang XuICML 2023 · 56 citations
- Information-theoretic Generalization Analysis for Expected Calibration ErrorFutoshi Futami, Masahiro FujisawaNeurIPS 2024 · 22 citations
- Improving Neural Additive Models with Bayesian PrinciplesKouroche Bouchiat, Alexander Immer, Hugo Yèche, Gunnar Rätsch et al.ICML 2024 · 17 citations
- Beyond probability partitions: Calibrating neural networks with semantic aware groupingJia-Qi Yang, De-Chuan Zhan, Le GanNeurIPS 2023 · 14 citations
- LaSCal: Label-Shift Calibration without target labelsTeodora Popordanoska, Gorjan Radevski, Tinne Tuytelaars, Matthew B. BlaschkoNeurIPS 2024 · 12 citations
Builds on17
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 344 citations
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 276 citations
- Be Confident! Towards Trustworthy Graph Neural Networks via Confidence CalibrationXiao Wang, Hongrui Liu, Chuan Shi, Cheng YangNeurIPS 2021 · 158 citations
- Beyond Pinball Loss: Quantile Methods for Calibrated Uncertainty QuantificationYoungseog Chung, Willie Neiswanger, Ian Char, Jeff SchneiderNeurIPS 2021 · 137 citations
Related papers
- Reassessing How to Compare and Improve the Calibration of Machine Learning ModelsMuthu Chidambaram, Rong GeICLR 2025
- PAC-Bayes Analysis for Recalibration in ClassificationMasahiro Fujisawa, Futoshi FutamiICML 2025
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 177 citations
- Beyond calibration: estimating the grouping loss of modern neural networksAlexandre Perez-Lebel, Marine Le Morvan, Gaël VaroquauxICLR 2023 · 5 citations
- Improving Perturbation-based Explanations by Understanding the Role of Uncertainty CalibrationThomas Decker, Volker Tresp, Florian BuettnerNeurIPS 2025 · 3 citations
