Never mind the metrics - what about the uncertainty? Visualising binary confusion matrix metric distributions to put performance in perspective
David R. Lovell, Dimity Miller, Jaiden Capra, Andrew P. Bradley
Abstract
There are strong incentives to build models that demonstrate outstanding predictive performance on various datasets and benchmarks. We believe these incentives risk a narrow focus on models and on the performance metrics used to evaluate and compare them-resulting in a growing body of literature to evaluate and compare metrics. This paper strives for a more balanced perspective on classifier performance metrics by highlighting their distributions under different models of uncertainty and showing how this uncertainty can easily eclipse differences in the empirical performance of classifiers. We begin by emphasising the fundamentally discrete nature of empirical confusion matrices and show how binary matrices can be meaningfully represented in a three dimensional compositional lattice, whose cross-sections form the basis of the space of receiver operating characteristic (ROC) curves. We develop equations, animations and interactive visualisations of the contours of performance metrics within (and beyond) this ROC space, showing how some are affected by class imbalance. We provide interactive visualisations that show the discrete posterior predictive probability mass functions of true and false positive rates in ROC space, and how these relate to uncertainty in performance metrics such as Balanced Accuracy (BA) and the Matthews Correlation Coefficient (MCC). Our hope is that these insights and visualisations will raise greater awareness of the substantial uncertainty in performance metric estimates that can arise when classifiers are evaluated on empirical datasets and benchmarks, and that classification model performance claims should be tempered by this understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03316dd1-9ce3-45b1-b15e-134ac091e542Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- The VOROS: Lifting ROC Curves to 3D to Summarize Unbalanced Classifier PerformanceChristopher Ratigan, Lenore CowenAAAI 2025 · 1 citation
- A Closer Look at AUROC and AUPRC under Class ImbalanceMatthew B. A. McDermott, Haoran Zhang, Lasse Hyldig Hansen, Giovanni Angelotti et al.NeurIPS 2024 · 191 citations
- Calibrate: Interactive Analysis of Probabilistic Model OutputPeter Xenopoulos, João Rulff, Luis Gustavo Nonato, Brian Barr et al.IEEE VIS 2022 · 19 citations
- Neo: Generalizing Confusion Matrix Visualization to Hierarchical and Multi-Output LabelsJochen Görtler, Fred Hohman, Dominik Moritz, Kanit Wongsuphasawat et al.CHI 2022 · 70 citations
- Interplay of ROC and Precision-Recall AUCs: Theoretical Limits and Practical Implications in Binary ClassificationMartin Mihelich, François Castagnos, Charles DogninICML 2024
