Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED
Quzhe Huang, Shibo Hao, Yuan Ye, Shengqi Zhu, Yansong Feng, Dongyan Zhao
Abstract
DocRED is a widely used dataset for document-level relation extraction. In the large-scale annotation, a recommend-revise scheme is adopted to reduce the workload. Within this scheme, annotators are provided with candidate relation instances from distant supervision, and they then manually supplement and remove relational facts based on the recommendations. However, when comparing DocRED with a subset relabeled from scratch, we find that this scheme results in a considerable amount of false negative samples and an obvious bias towards popular entities and relations. Furthermore, we observe that the models trained on DocRED have low recall on our relabeled dataset and inherit the same bias in the training data. Through the analysis of annotators' behaviors, we figure out the underlying reason for the problems above: the scheme actually discourages annotators from supplementing adequate instances in the revision phase. We appeal to future research to take into consideration the issues with the recommend-revise scheme when designing new models and annotation schemes. The relabeled dataset is released at https://github.com/AndrewZhe/ Revisit-DocRED , to serve as a more reliable test set of document RE models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f4f6e1a-5ae7-4c47-b6c9-baf7d3239ccdCited by top-tier papers11
- Revisiting DocRED - Addressing the False Negative Problem in Relation ExtractionQingyu Tan, Lu Xu, Lidong Bing, Hwee Tou Ng et al.EMNLP 2022 · 76 citations
- A Unified Positive-Unlabeled Learning Framework for Document-Level Relation Extraction with Different Levels of LabelingYe Wang, Xinxin Liu, Wenxin Hu, Tao ZhangEMNLP 2022 · 18 citations
- Anaphor Assisted Document-Level Relation ExtractionChonggang Lu, Richong Zhang, Kai Sun, Jaein Kim et al.EMNLP 2023 · 16 citations
- A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete LabelingYe Wang, Huazheng Pan, Tao Zhang, Wen Wu et al.AAAI 2024 · 11 citations
- Boosting Document-Level Relation Extraction by Mining and Injecting Logical RulesShengda Fan, Shasha Mo, Jianwei NiuEMNLP 2022 · 9 citations
Builds on7
- Document-Level Relation Extraction with Adaptive Thresholding and Localized Context PoolingWenxuan Zhou, Kevin Huang, Tengyu Ma, Jing HuangAAAI 2021 · 360 citations
- Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionGuoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei LuACL 2020 · 294 citations
- Double Graph Based Reasoning for Document-level Relation ExtractionShuang Zeng, Runxin Xu, Baobao Chang, Lei LiEMNLP 2020 · 238 citations
- Entity Structure Within and Throughout: Modeling Mention Dependencies for Document-Level Relation ExtractionBenfeng Xu, Quan Wang, Yajuan Lyu, Yong Zhu et al.AAAI 2021 · 200 citations
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 146 citations
Related papers
- TTM-RE: Memory-Augmented Document-Level Relation ExtractionChufan Gao, Xuan Wang, Jimeng SunACL 2024 · 11 citations
- TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction TaskChristoph Alt, Aleksandra Gabryszak, Leonhard HennigACL 2020 · 9 citations
- SENT: Sentence-level Distant Relation Extraction via Negative TrainingRuotian Ma, Tao Gui, Linyang Li, Qi Zhang et al.ACL 2021
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li et al.EMNLP 2021 · 18 citations
- Revisiting the Negative Data of Distantly Supervised Relation ExtractionChenhao Xie, Jiaqing Liang, Jingping Liu, Chengsong Huang et al.ACL 2021
