TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction Task
Christoph Alt, Aleksandra Gabryszak, Leonhard Hennig
摘要
TACRED (Zhang et al., 2017) is one of the largest, most widely used crowdsourced datasets in Relation Extraction (RE). But, even with recent advances in unsupervised pretraining and knowledge enhanced neural RE, models still show a high error rate. In this paper, we investigate the questions: Have we reached a performance ceiling or is there still room for improvement? And how do crowd annotations, dataset, and models contribute to this error rate? To answer these questions, we first validate the most challenging 5K examples in the development and test sets using trained annotators. We find that label errors account for 8% absolute F1 test error, and that more than 50% of the examples need to be relabeled. On the relabeled test set the average F1 score of a large baseline model set improves from 62.1 to 70.1. After validation, we analyze misclassifications on the challenging instances, categorize them into linguistically motivated error groups, and verify the resulting error hypotheses on three state-of-the-art RE models. We show that two groups of ambiguous relations are responsible for most of the remaining errors and that models may adopt shallow heuristics on the dataset when entities are not masked.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation ExtractionXiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng 等WWW 2022 · 被引用 488 次
- Differentiable Prompt Makes Pre-trained Language Models Better Few-shot LearnersNingyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng 等ICLR 2022 · 被引用 205 次
- Learning from Context or Names? An Empirical Study on Neural Relation ExtractionHao Peng, Tianyu Gao, Xu Han, Yankai Lin 等EMNLP 2020 · 被引用 185 次
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 被引用 146 次
- Label Verbalization and Entailment for Effective Zero and Few-Shot Relation ExtractionOscar Sainz, Oier Lopez de Lacalle, Gorka Labaka, Ander Barrena 等EMNLP 2021 · 被引用 94 次
相关 Paper
- Revisiting DocRED - Addressing the False Negative Problem in Relation ExtractionQingyu Tan, Lu Xu, Lidong Bing, Hwee Tou Ng 等EMNLP 2022 · 被引用 76 次
- MultiTACRED: A Multilingual Version of the TAC Relation Extraction DatasetLeonhard Hennig, Philippe Thomas, Sebastian MöllerACL 2023 · 被引用 4 次
- Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocREDQuzhe Huang, Shibo Hao, Yuan Ye, Shengqi Zhu 等ACL 2022
- Probing Linguistic Features of Sentence-Level Representations in Relation ExtractionChristoph Alt, Aleksandra Gabryszak, Leonhard HennigACL 2020 · 被引用 29 次
- TTM-RE: Memory-Augmented Document-Level Relation ExtractionChufan Gao, Xuan Wang, Jimeng SunACL 2024 · 被引用 11 次
