An Effectiveness Metric for Ordinal Classification: Formal Properties and Experimental Results
Enrique Amigó, Julio Gonzalo, Stefano Mizzaro, Jorge Carrillo-de-Albornoz
摘要
In Ordinal Classification tasks, items have to be assigned to classes that have a relative ordering, such as positive, neutral, negative in sentiment analysis. Remarkably, the most popular evaluation metrics for ordinal classification tasks either ignore relevant information (for instance, precision/recall on each of the classes ignores their relative ordering) or assume additional information (for instance, Mean Average Error assumes absolute distances between classes). In this paper we propose a new metric for Ordinal Classification, Closeness Evaluation Measure, that is rooted on Measurement Theory and Information Theory. Our theoretical analysis and experimental results over both synthetic data and data from NLP shared tasks indicate that the proposed metric captures quality aspects from different traditional tasks simultaneously. In addition, it generalizes some popular classification (nominal scale) and error minimization (interval scale) metrics, depending on the measurement scale in which it is instantiated.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Ranking Interruptus: When Truncated Rankings Are Better and How to Measure ThatEnrique Amigó, Stefano Mizzaro, Damiano SpinaSIGIR 2022 · 被引用 6 次
- SLACE: A Monotone and Balance-Sensitive Loss Function for Ordinal RegressionInbar Nachmani, Bar Genossar, Coral Scharf, Roee Shraga 等AAAI 2025 · 被引用 1 次
- Evaluating Evaluation Measures for Ordinal Classification and Ordinal QuantificationTetsuya SakaiACL 2021
- Contrastive Order Learning: A General Framework for Ordinal RegressionChaewon Lee, BeomJun Shim, Kwang Choi, Chang-Su KimICML 2026
- A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better InterpretabilityXinyu Hu, Mingqi Gao, Li Lin, Zhenghan Yu 等ACL 2025
相关 Paper
- Evaluating Extreme Hierarchical Multi-label ClassificationEnrique Amigó, Agustín D. DelgadoACL 2022
- To Err Like Human: Affective Bias-Inspired Measures for Visual Emotion Recognition EvaluationChenxi Zhao, Jinglei Shi, Liqiang Nie, Jufeng YangNeurIPS 2024 · 被引用 10 次
- An Ordinal Data Clustering Algorithm with Automated Distance LearningYiqun Zhang, Yiu-ming CheungAAAI 2020 · 被引用 28 次
- Better than Average: Paired Evaluation of NLP systemsMaxime Peyrard, Wei Zhao, Steffen Eger, Robert WestACL 2021
- Field-aware Calibration: A Simple and Empirically Strong Method for Reliable Probabilistic PredictionsFeiyang Pan, Xiang Ao, Pingzhong Tang, Min Lu 等WWW 2020 · 被引用 30 次
