How Flawed Is ECE? An Analysis via Logit Smoothing
Muthu Chidambaram, Holden Lee, Colin McSwiggen, Semon Rezchikov
Abstract
Informally, a model is calibrated if its predictions are correct with a probability that matches the confidence of the prediction. By far the most common method in the literature for measuring calibration is the expected calibration error (ECE). Recent work, however, has pointed out drawbacks of ECE, such as the fact that it is discontinuous in the space of predictors. In this work, we ask: how fundamental are these issues, and what are their impacts on existing results? Towards this end, we completely characterize the discontinuities of ECE with respect to general probability measures on Polish spaces. We then use the nature of these discontinuities to motivate a novel continuous, easily estimated miscalibration metric, which we term Logit-Smoothed ECE (LS-ECE). By comparing the ECE and LS-ECE of pre-trained image classification models, we show in initial experiments that binned ECE closely tracks LS-ECE, indicating that the theoretical pathologies of ECE may be avoidable in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ba14e9a-a084-4cb1-b677-647c3a1b6db5Cited by top-tier papers3
- Combining Priors with Experience: Confidence Calibration Based on Binomial Process ModelingJinzong Dong, Zhaohui Jiang, Dong Pan, Haoyang YuAAAI 2025 · 4 citations
- When High Accuracy Hides Poor Calibration: Rethinking Confidence Evaluation in Transformer-Based Text Classification with Balanced Brier ScoreGuilherme Fonseca, Gabriel Prenassi, Washington Cunha, Leonardo Chaves Dutra da Rocha et al.ACL 2026
- Reassessing How to Compare and Improve the Calibration of Machine Learning ModelsMuthu Chidambaram, Rong GeICLR 2025
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 276 citations
Related papers
- A Unifying Theory of Distance from CalibrationJaroslaw Blasiok, Parikshit Gopalan, Lunjia Hu, Preetum NakkiranSTOC 2023 · 7 citations
- Smooth ECE: Principled Reliability Diagrams via Kernel SmoothingJaroslaw Blasiok, Preetum NakkiranICLR 2024 · 59 citations
- Variable-Based Calibration for Machine Learning ClassifiersMarkelle Kelly, Padhraic SmythAAAI 2023 · 7 citations
- Information-theoretic Generalization Analysis for Expected Calibration ErrorFutoshi Futami, Masahiro FujisawaNeurIPS 2024 · 22 citations
- An Elementary Predictor Obtaining Distance to CalibrationEshwar Ram Arunachaleswaran, Natalie Collina, Aaron Roth, Mirah ShiSODA 2025
