Never mind the metrics - what about the uncertainty? Visualising binary confusion matrix metric distributions to put performance in perspective
David R. Lovell, Dimity Miller, Jaiden Capra, Andrew P. Bradley
摘要
There are strong incentives to build models that demonstrate outstanding predictive performance on various datasets and benchmarks. We believe these incentives risk a narrow focus on models and on the performance metrics used to evaluate and compare them-resulting in a growing body of literature to evaluate and compare metrics. This paper strives for a more balanced perspective on classifier performance metrics by highlighting their distributions under different models of uncertainty and showing how this uncertainty can easily eclipse differences in the empirical performance of classifiers. We begin by emphasising the fundamentally discrete nature of empirical confusion matrices and show how binary matrices can be meaningfully represented in a three dimensional compositional lattice, whose cross-sections form the basis of the space of receiver operating characteristic (ROC) curves. We develop equations, animations and interactive visualisations of the contours of performance metrics within (and beyond) this ROC space, showing how some are affected by class imbalance. We provide interactive visualisations that show the discrete posterior predictive probability mass functions of true and false positive rates in ROC space, and how these relate to uncertainty in performance metrics such as Balanced Accuracy (BA) and the Matthews Correlation Coefficient (MCC). Our hope is that these insights and visualisations will raise greater awareness of the substantial uncertainty in performance metric estimates that can arise when classifiers are evaluated on empirical datasets and benchmarks, and that classification model performance claims should be tempered by this understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- The VOROS: Lifting ROC Curves to 3D to Summarize Unbalanced Classifier PerformanceChristopher Ratigan, Lenore CowenAAAI 2025 · 被引用 1 次
- A Closer Look at AUROC and AUPRC under Class ImbalanceMatthew B. A. McDermott, Haoran Zhang, Lasse Hyldig Hansen, Giovanni Angelotti 等NeurIPS 2024 · 被引用 191 次
- Calibrate: Interactive Analysis of Probabilistic Model OutputPeter Xenopoulos, João Rulff, Luis Gustavo Nonato, Brian Barr 等IEEE VIS 2022 · 被引用 19 次
- Neo: Generalizing Confusion Matrix Visualization to Hierarchical and Multi-Output LabelsJochen Görtler, Fred Hohman, Dominik Moritz, Kanit Wongsuphasawat 等CHI 2022 · 被引用 70 次
- Interplay of ROC and Precision-Recall AUCs: Theoretical Limits and Practical Implications in Binary ClassificationMartin Mihelich, François Castagnos, Charles DogninICML 2024
