Evaluating Sequence Labeling on the basis of Information Theory
Enrique Amigó, Elena Álvarez Mellado, Julio Gonzalo, Jorge Carrillo-de-Albornoz
Abstract
Various metrics exist for evaluating sequence labeling problems (strict span matching, token oriented metrics, token concurrence in sequences, etc.), each of them focusing on certain aspects of the task. In this paper, we define a comprehensive set of formal properties that captures the strengths and weaknesses of the existing metric families and prove that none of them is able to satisfy all properties simultaneously. We argue that it is necessary to measure how much information (correct or noisy) each token in the sequence contributes depending on different aspects such as sequence length, number of tokens annotated by the system, token specificity, etc. On this basis, we introduce the Sequence Labelling Information Contrast Model (SL-ICM), a novel metric based on information theory for evaluating sequence labeling tasks. Our formal analysis and experimentation show that the proposed metric satisfies all properties simultaneously.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
Related papers
- Estimating Agreement by Chance for Sequence AnnotationDiya Li, Carolyn P. Rosé, Ao Yuan, Chunxiao ZhouACL 2024
- Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language GenerationZhen Lin, Shubhendu Trivedi, Jimeng SunEMNLP 2024 · 1 citation
- Neural-Hidden-CRF: A Robust Weakly-Supervised Sequence LabelerZhijun Chen, Hailong Sun, Wanhao Zhang, Chunyi Xu et al.KDD 2023 · 1 citation
- Solving Sequential Text Classification as Board-Game PlayingChen Qian, Fuli Feng, Lijie Wen, Zhenpeng Chen et al.AAAI 2020 · 2 citations
- AvgOut: A Simple Output-Probability Measure to Eliminate Dull ResponsesTong Niu, Mohit BansalAAAI 2020 · 3 citations
