How Flawed Is ECE? An Analysis via Logit Smoothing
Muthu Chidambaram, Holden Lee, Colin McSwiggen, Semon Rezchikov
摘要
Informally, a model is calibrated if its predictions are correct with a probability that matches the confidence of the prediction. By far the most common method in the literature for measuring calibration is the expected calibration error (ECE). Recent work, however, has pointed out drawbacks of ECE, such as the fact that it is discontinuous in the space of predictors. In this work, we ask: how fundamental are these issues, and what are their impacts on existing results? Towards this end, we completely characterize the discontinuities of ECE with respect to general probability measures on Polish spaces. We then use the nature of these discontinuities to motivate a novel continuous, easily estimated miscalibration metric, which we term Logit-Smoothed ECE (LS-ECE). By comparing the ECE and LS-ECE of pre-trained image classification models, we show in initial experiments that binned ECE closely tracks LS-ECE, indicating that the theoretical pathologies of ECE may be avoidable in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Combining Priors with Experience: Confidence Calibration Based on Binomial Process ModelingJinzong Dong, Zhaohui Jiang, Dong Pan, Haoyang YuAAAI 2025 · 被引用 4 次
- When High Accuracy Hides Poor Calibration: Rethinking Confidence Evaluation in Transformer-Based Text Classification with Balanced Brier ScoreGuilherme Fonseca, Gabriel Prenassi, Washington Cunha, Leonardo Chaves Dutra da Rocha 等ACL 2026
- Reassessing How to Compare and Improve the Calibration of Machine Learning ModelsMuthu Chidambaram, Rong GeICLR 2025
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis 等NeurIPS 2021 · 被引用 633 次
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 被引用 276 次
相关 Paper
- A Unifying Theory of Distance from CalibrationJaroslaw Blasiok, Parikshit Gopalan, Lunjia Hu, Preetum NakkiranSTOC 2023 · 被引用 7 次
- Smooth ECE: Principled Reliability Diagrams via Kernel SmoothingJaroslaw Blasiok, Preetum NakkiranICLR 2024 · 被引用 59 次
- Variable-Based Calibration for Machine Learning ClassifiersMarkelle Kelly, Padhraic SmythAAAI 2023 · 被引用 7 次
- Information-theoretic Generalization Analysis for Expected Calibration ErrorFutoshi Futami, Masahiro FujisawaNeurIPS 2024 · 被引用 22 次
- An Elementary Predictor Obtaining Distance to CalibrationEshwar Ram Arunachaleswaran, Natalie Collina, Aaron Roth, Mirah ShiSODA 2025
