Good Classification Measures and How to Find Them
Martijn Gösgens, Anton Zhiyanov, Aleksey Tikhonov, Liudmila Prokhorenkova
Abstract
Several performance measures can be used for evaluating classification results: accuracy, F-measure, and many others. Can we say that some of them are better than others, or, ideally, choose one measure that is best in all situations? To answer this question, we conduct a systematic analysis of classification performance measures: we formally define a list of desirable properties and theoretically analyze which measures satisfy which properties. We also prove an impossibility theorem: some desirable properties cannot be simultaneously satisfied. Finally, we propose a new family of measures satisfying all desirable properties except one. This family includes the Matthews Correlation Coefficient and a so-called Symmetric Balanced Accuracy that was not previously used in classification literature. We believe that our systematic approach gives an important tool to practitioners for adequately evaluating classification results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f762fbd2-272f-4f4d-bd00-be72a1e4d1eeCited by top-tier papers2
- Characterizing Graph Datasets for Node Classification: Homophily-Heterophily Dichotomy and BeyondOleg Platonov, Denis Kuznedelev, Artem Babenko, Liudmila ProkhorenkovaNeurIPS 2023 · 95 citations
- Never mind the metrics - what about the uncertainty? Visualising binary confusion matrix metric distributions to put performance in perspectiveDavid R. Lovell, Dimity Miller, Jaiden Capra, Andrew P. BradleyICML 2023 · 3 citations
Builds on3
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Systematic Analysis of Cluster Similarity Indices: How to Validate Validation MeasuresMartijn Gösgens, Alexey Tikhonov, Liudmila ProkhorenkovaICML 2021 · 27 citations
- Self-Training With Noisy Student Improves ImageNet ClassificationQizhe Xie, Minh-Thang Luong, Eduard H. Hovy, Quoc V. LeCVPR 2020
Related papers
- What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is Rarely F1Sébastien Piérard, Adrien Deliege, Marc Van DroogenbroeckCVPR 2026 · 2 citations
- On the use of evaluation measures for defect prediction studiesRebecca Moussa, Federica SarroISSTA 2022 · 39 citations
- Foundations of the Theory of Performance-Based RankingSébastien Piérard, Anaïs Halin, Anthony Cioppa, Adrien Deliège et al.CVPR 2025
- Optimal Decision Trees for Nonlinear MetricsEmir Demirovic, Peter J. StuckeyAAAI 2021 · 29 citations
- Simple Weak Coresets for Non-decomposable Classification MeasuresJayesh Malaviya, Anirban Dasgupta, Rachit ChhayaAAAI 2024 · 1 citation
