Reassessing How to Compare and Improve the Calibration of Machine Learning Models
Muthu Chidambaram, Rong Ge
摘要
A machine learning model is calibrated if its predicted probability for an outcome matches the observed frequency for that outcome conditional on the model prediction. This property has become increasingly important as the impact of machine learning models has continued to spread to various domains. As a result, there are now a dizzying number of recent papers on measuring and improving the calibration of (specifically deep learning) models. In this work, we reassess the reporting of calibration metrics in the recent literature. We show that there exist trivial recalibration approaches that can appear seemingly state-of-the-art unless calibration and prediction metrics (i.e. test accuracy) are accompanied by additional generalization metrics such as negative log-likelihood. We then use a calibration-based decomposition of Bregman divergences to develop a new extension to reliability diagrams that jointly visualizes calibration and generalization error, and show how our visualization can be used to detect trade-offs between calibration and generalization. Along the way, we prove novel results regarding the relationship between full calibration error and confidence calibration error for Bregman divergences. We also establish the consistency of the kernel regression estimator for calibration error used in our visualization approach, which generalizes existing consistency results in the literature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis 等NeurIPS 2021 · 被引用 633 次
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 被引用 276 次
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 被引用 177 次
相关 Paper
- Calibrate: Interactive Analysis of Probabilistic Model OutputPeter Xenopoulos, João Rulff, Luis Gustavo Nonato, Brian Barr 等IEEE VIS 2022 · 被引用 19 次
- Smooth ECE: Principled Reliability Diagrams via Kernel SmoothingJaroslaw Blasiok, Preetum NakkiranICLR 2024 · 被引用 59 次
- Calibration tests beyond classificationDavid Widmann, Fredrik Lindsten, Dave ZachariahICLR 2021 · 被引用 23 次
- Better Uncertainty Calibration via Proper Scores for Classification and BeyondSebastian G. Gruber, Florian BuettnerNeurIPS 2022 · 被引用 88 次
- Calibrated and Sharp Uncertainties in Deep Learning via Density EstimationVolodymyr Kuleshov, Shachi DeshpandeICML 2022 · 被引用 44 次
