Calibration tests beyond classification
David Widmann, Fredrik Lindsten, Dave Zachariah
摘要
Most supervised machine learning tasks are subject to irreducible prediction errors. Probabilistic predictive models address this limitation by providing probability distributions that represent a belief over plausible targets, rather than point estimates. Such models can be a valuable tool in decision-making under uncertainty, provided that the model output is meaningful and interpretable. Calibrated models guarantee that the probabilistic predictions are neither over-nor under-confident. In the machine learning literature, different measures and statistical tests have been proposed and studied for evaluating the calibration of classification models. For regression problems, however, research has been focused on a weaker condition of calibration based on predicted quantiles for real-valued targets. In this paper, we propose the first framework that unifies calibration evaluation and tests for general probabilistic predictive models. It applies to any such model, including classification and regression models of arbitrary dimension. Furthermore, the framework generalizes existing measures and provides a more intuitive reformulation of a recently proposed framework for calibration in multi-class classification. In particular, we reformulate and generalize the kernel calibration error, its estimators, and hypothesis tests using scalar-valued kernels, and evaluate the calibration of real-valued regression problems. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Better Uncertainty Calibration via Proper Scores for Classification and BeyondSebastian G. Gruber, Florian BuettnerNeurIPS 2022 · 被引用 88 次
- Smooth ECE: Principled Reliability Diagrams via Kernel SmoothingJaroslaw Blasiok, Preetum NakkiranICLR 2024 · 被引用 59 次
- Information-theoretic Generalization Analysis for Expected Calibration ErrorFutoshi Futami, Masahiro FujisawaNeurIPS 2024 · 被引用 22 次
- Calibration by Distribution Matching: Trainable Kernel Calibration MetricsCharlie Marx, Sofian Zalouk, Stefano ErmonNeurIPS 2023 · 被引用 21 次
- Improving Neural Additive Models with Bayesian PrinciplesKouroche Bouchiat, Alexander Immer, Hugo Yèche, Gunnar Rätsch 等ICML 2024 · 被引用 17 次
它引用的顶会 Paper1
相关 Paper
- A Unifying Theory of Distance from CalibrationJaroslaw Blasiok, Parikshit Gopalan, Lunjia Hu, Preetum NakkiranSTOC 2023 · 被引用 7 次
- PAC-Bayes Analysis for Recalibration in ClassificationMasahiro Fujisawa, Futoshi FutamiICML 2025
- Reassessing How to Compare and Improve the Calibration of Machine Learning ModelsMuthu Chidambaram, Rong GeICLR 2025
- Nonparametric Distribution Regression Re-calibrationÁdám Jung, Domokos Kelen, Andras BenczurICML 2026
- Field-aware Calibration: A Simple and Empirically Strong Method for Reliable Probabilistic PredictionsFeiyang Pan, Xiang Ao, Pingzhong Tang, Min Lu 等WWW 2020 · 被引用 30 次
