Re-TACRED: Addressing Shortcomings of the TACRED Dataset
George Stoica, Emmanouil Antonios Platanios, Barnabás Póczos
Abstract
TACRED is one of the largest and most widely used sentence-level relation extraction datasets. Proposed models that are evaluated using this dataset consistently set new state-of-the-art performance. However, they still exhibit large error rates despite leveraging external knowledge and unsupervised pretraining on large text corpora. A recent study suggested that this may be due to poor dataset quality. The study observed that over 50% of the most challenging sentences from the development and test sets are incorrectly labeled and account for an average drop of 8% f1-score in model performance. However, this study was limited to a small biased sample of 5k (out of a total of 106k) sentences, substantially restricting the generalizability and broader implications of its findings. In this paper, we address these shortcomings by: (i) performing a comprehensive study over the whole TACRED dataset, (ii) proposing an improved crowdsourcing strategy and deploying it to re-annotate the whole dataset, and (iii) performing a thorough analysis to understand how correcting the TACRED annotations affects previously published results. After verification, we observed that 23.9% of TACRED labels are incorrect. Moreover, evaluating several models on our revised dataset yields an average f1-score improvement of 14.3% and helps uncover significant relationships between the different models (rather than simply offsetting or scaling their scores by a constant factor). Finally, aside from our analysis we also release Re-TACRED, a new completely re-annotated version of the TACRED dataset that can be used to perform reliable evaluation of relation extraction models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5338f48-425d-4f2b-b879-3e9d1cff8b6aCited by top-tier papers17
- S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation ExtractionBenfeng Xu, Quan Wang, Yajuan Lyu, Dai Dai et al.ACL 2023 · 13 citations
- BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain AbstractionJiangmeng Li, Fei Song, Yifan Jin, Wenwen Qiang et al.ICLR 2024 · 9 citations
- HistRED: A Historical Document-Level Relation Extraction DatasetSoyoung Yang, Minseok Choi, Youngwoo Cho, Jaegul ChooACL 2023 · 9 citations
- REDFM: a Filtered and Multilingual Relation Extraction DatasetPere-Lluís Huguet Cabot, Simone Tedeschi, Axel-Cyrille Ngonga Ngomo, Roberto NavigliACL 2023 · 9 citations
- Revisiting the Knowledge Injection FrameworksPeng Fu, Yiming Zhang, Haobo Wang, Weikang Qiu et al.EMNLP 2023 · 8 citations
Builds on1
Related papers
- Revisiting DocRED - Addressing the False Negative Problem in Relation ExtractionQingyu Tan, Lu Xu, Lidong Bing, Hwee Tou Ng et al.EMNLP 2022 · 76 citations
- Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocREDQuzhe Huang, Shibo Hao, Yuan Ye, Shengqi Zhu et al.ACL 2022
- MultiTACRED: A Multilingual Version of the TAC Relation Extraction DatasetLeonhard Hennig, Philippe Thomas, Sebastian MöllerACL 2023 · 4 citations
- TTM-RE: Memory-Augmented Document-Level Relation ExtractionChufan Gao, Xuan Wang, Jimeng SunACL 2024 · 11 citations
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li et al.EMNLP 2021 · 18 citations
