Re-TACRED: Addressing Shortcomings of the TACRED Dataset
George Stoica, Emmanouil Antonios Platanios, Barnabás Póczos
摘要
TACRED is one of the largest and most widely used sentence-level relation extraction datasets. Proposed models that are evaluated using this dataset consistently set new state-of-the-art performance. However, they still exhibit large error rates despite leveraging external knowledge and unsupervised pretraining on large text corpora. A recent study suggested that this may be due to poor dataset quality. The study observed that over 50% of the most challenging sentences from the development and test sets are incorrectly labeled and account for an average drop of 8% f1-score in model performance. However, this study was limited to a small biased sample of 5k (out of a total of 106k) sentences, substantially restricting the generalizability and broader implications of its findings. In this paper, we address these shortcomings by: (i) performing a comprehensive study over the whole TACRED dataset, (ii) proposing an improved crowdsourcing strategy and deploying it to re-annotate the whole dataset, and (iii) performing a thorough analysis to understand how correcting the TACRED annotations affects previously published results. After verification, we observed that 23.9% of TACRED labels are incorrect. Moreover, evaluating several models on our revised dataset yields an average f1-score improvement of 14.3% and helps uncover significant relationships between the different models (rather than simply offsetting or scaling their scores by a constant factor). Finally, aside from our analysis we also release Re-TACRED, a new completely re-annotated version of the TACRED dataset that can be used to perform reliable evaluation of relation extraction models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation ExtractionBenfeng Xu, Quan Wang, Yajuan Lyu, Dai Dai 等ACL 2023 · 被引用 13 次
- BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain AbstractionJiangmeng Li, Fei Song, Yifan Jin, Wenwen Qiang 等ICLR 2024 · 被引用 9 次
- HistRED: A Historical Document-Level Relation Extraction DatasetSoyoung Yang, Minseok Choi, Youngwoo Cho, Jaegul ChooACL 2023 · 被引用 9 次
- REDFM: a Filtered and Multilingual Relation Extraction DatasetPere-Lluís Huguet Cabot, Simone Tedeschi, Axel-Cyrille Ngonga Ngomo, Roberto NavigliACL 2023 · 被引用 9 次
- Revisiting the Knowledge Injection FrameworksPeng Fu, Yiming Zhang, Haobo Wang, Weikang Qiu 等EMNLP 2023 · 被引用 8 次
它引用的顶会 Paper1
相关 Paper
- Revisiting DocRED - Addressing the False Negative Problem in Relation ExtractionQingyu Tan, Lu Xu, Lidong Bing, Hwee Tou Ng 等EMNLP 2022 · 被引用 76 次
- Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocREDQuzhe Huang, Shibo Hao, Yuan Ye, Shengqi Zhu 等ACL 2022
- MultiTACRED: A Multilingual Version of the TAC Relation Extraction DatasetLeonhard Hennig, Philippe Thomas, Sebastian MöllerACL 2023 · 被引用 4 次
- TTM-RE: Memory-Augmented Document-Level Relation ExtractionChufan Gao, Xuan Wang, Jimeng SunACL 2024 · 被引用 11 次
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li 等EMNLP 2021 · 被引用 18 次
