Calibrate: Interactive Analysis of Probabilistic Model Output
Peter Xenopoulos, João Rulff, Luis Gustavo Nonato, Brian Barr, Cláudio T. Silva
Abstract
Analyzing classification model performance is a crucial task for machine learning practitioners. While practitioners often use count-based metrics derived from confusion matrices, like accuracy, many applications, such as weather prediction, sports betting, or patient risk prediction, rely on a classifier's predicted probabilities rather than predicted labels. In these instances, practitioners are concerned with producing a calibrated model, that is, one which outputs probabilities that reflect those of the true distribution. Model calibration is often analyzed visually, through static reliability diagrams, however, the traditional calibration visualization may suffer from a variety of drawbacks due to the strong aggregations it necessitates. Furthermore, count-based approaches are unable to sufficiently analyze model calibration. We present Calibrate, an interactive reliability diagram that addresses the aforementioned issues. Calibrate constructs a reliability diagram that is resistant to drawbacks in traditional approaches, and allows for interactive subgroup analysis and instance-level inspection. We demonstrate the utility of Calibrate through use cases on both real-world and synthetic data. We further validate Calibrate by presenting the results of a think-aloud experiment with data scientists who routinely analyze model calibration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fd37f35-ede3-4f74-b39f-a03f41873020Cited by top-tier papers3
- Notable: On-the-fly Assistant for Data Storytelling in Computational NotebooksHaotian Li, Lu Ying, Haidong Zhang, Yingcai Wu et al.CHI 2023 · 38 citations
- A Unified Interactive Model Evaluation for Classification, Object Detection, and Instance Segmentation in Computer VisionChangjian Chen, Yukai Guo, Fengyuan Tian, Shilong Liu et al.IEEE VIS 2023 · 28 citations
- Crowdsourced Think-Aloud StudiesZach Cutler, Lane Harrison, Carolina Nobre, Alexander LexCHI 2025 · 4 citations
Builds on1
Related papers
- Reassessing How to Compare and Improve the Calibration of Machine Learning ModelsMuthu Chidambaram, Rong GeICLR 2025
- Smooth ECE: Principled Reliability Diagrams via Kernel SmoothingJaroslaw Blasiok, Preetum NakkiranICLR 2024 · 59 citations
- Never mind the metrics - what about the uncertainty? Visualising binary confusion matrix metric distributions to put performance in perspectiveDavid R. Lovell, Dimity Miller, Jaiden Capra, Andrew P. BradleyICML 2023 · 3 citations
- Variable-Based Calibration for Machine Learning ClassifiersMarkelle Kelly, Padhraic SmythAAAI 2023 · 7 citations
- Calibration tests beyond classificationDavid Widmann, Fredrik Lindsten, Dave ZachariahICLR 2021 · 23 citations
