Good Classification Measures and How to Find Them
Martijn Gösgens, Anton Zhiyanov, Aleksey Tikhonov, Liudmila Prokhorenkova
摘要
Several performance measures can be used for evaluating classification results: accuracy, F-measure, and many others. Can we say that some of them are better than others, or, ideally, choose one measure that is best in all situations? To answer this question, we conduct a systematic analysis of classification performance measures: we formally define a list of desirable properties and theoretically analyze which measures satisfy which properties. We also prove an impossibility theorem: some desirable properties cannot be simultaneously satisfied. Finally, we propose a new family of measures satisfying all desirable properties except one. This family includes the Matthews Correlation Coefficient and a so-called Symmetric Balanced Accuracy that was not previously used in classification literature. We believe that our systematic approach gives an important tool to practitioners for adequately evaluating classification results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Characterizing Graph Datasets for Node Classification: Homophily-Heterophily Dichotomy and BeyondOleg Platonov, Denis Kuznedelev, Artem Babenko, Liudmila ProkhorenkovaNeurIPS 2023 · 被引用 95 次
- Never mind the metrics - what about the uncertainty? Visualising binary confusion matrix metric distributions to put performance in perspectiveDavid R. Lovell, Dimity Miller, Jaiden Capra, Andrew P. BradleyICML 2023 · 被引用 3 次
它引用的顶会 Paper3
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Systematic Analysis of Cluster Similarity Indices: How to Validate Validation MeasuresMartijn Gösgens, Alexey Tikhonov, Liudmila ProkhorenkovaICML 2021 · 被引用 27 次
- Self-Training With Noisy Student Improves ImageNet ClassificationQizhe Xie, Minh-Thang Luong, Eduard H. Hovy, Quoc V. LeCVPR 2020
相关 Paper
- What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is Rarely F1Sébastien Piérard, Adrien Deliege, Marc Van DroogenbroeckCVPR 2026 · 被引用 2 次
- On the use of evaluation measures for defect prediction studiesRebecca Moussa, Federica SarroISSTA 2022 · 被引用 39 次
- Foundations of the Theory of Performance-Based RankingSébastien Piérard, Anaïs Halin, Anthony Cioppa, Adrien Deliège 等CVPR 2025
- Optimal Decision Trees for Nonlinear MetricsEmir Demirovic, Peter J. StuckeyAAAI 2021 · 被引用 29 次
- Simple Weak Coresets for Non-decomposable Classification MeasuresJayesh Malaviya, Anirban Dasgupta, Rachit ChhayaAAAI 2024 · 被引用 1 次
