Variable-Based Calibration for Machine Learning Classifiers
Markelle Kelly, Padhraic Smyth
摘要
The deployment of machine learning classifiers in high-stakes domains requires well-calibrated confidence scores for model predictions. In this paper we introduce the notion of variable-based calibration to characterize calibration properties of a model with respect to a variable of interest, generalizing traditional score-based metrics such as expected calibration error (ECE). In particular, we find that models with near-perfect ECE can exhibit significant miscalibration as a function of features of the data. We demonstrate this phenomenon both theoretically and in practice on multiple well-known datasets, and show that it can persist after the application of existing calibration methods. To mitigate this issue, we propose strategies for detection, visualization, and quantification of variable-based calibration error. We then examine the limitations of current score-based calibration methods and explore potential modifications. Finally, we discuss the implications of these findings, emphasizing that an understanding of calibration beyond simple aggregate measures is crucial for endeavors such as fairness and model interpretability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 被引用 362 次
- Is the Most Accurate AI the Best Teammate? Optimizing AI for TeamworkGagan Bansal, Besmira Nushi, Ece Kamar, Eric Horvitz 等AAAI 2021 · 被引用 185 次
- Classification Under Human AssistanceAbir De, Nastaran Okati, Ali Zarezade, Manuel Gomez RodriguezAAAI 2021 · 被引用 60 次
- Field-aware Calibration: A Simple and Empirically Strong Method for Reliable Probabilistic PredictionsFeiyang Pan, Xiang Ao, Pingzhong Tang, Min Lu 等WWW 2020 · 被引用 30 次
相关 Paper
- How Flawed Is ECE? An Analysis via Logit SmoothingMuthu Chidambaram, Holden Lee, Colin McSwiggen, Semon RezchikovICML 2024 · 被引用 7 次
- When High Accuracy Hides Poor Calibration: Rethinking Confidence Evaluation in Transformer-Based Text Classification with Balanced Brier ScoreGuilherme Fonseca, Gabriel Prenassi, Washington Cunha, Leonardo Chaves Dutra da Rocha 等ACL 2026
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 被引用 276 次
- Smooth ECE: Principled Reliability Diagrams via Kernel SmoothingJaroslaw Blasiok, Preetum NakkiranICLR 2024 · 被引用 59 次
- Proximity-Informed Calibration for Deep Neural NetworksMiao Xiong, Ailin Deng, Pang Wei Koh, Jiaying Wu 等NeurIPS 2023 · 被引用 30 次
