Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED
Quzhe Huang, Shibo Hao, Yuan Ye, Shengqi Zhu, Yansong Feng, Dongyan Zhao
摘要
DocRED is a widely used dataset for document-level relation extraction. In the large-scale annotation, a recommend-revise scheme is adopted to reduce the workload. Within this scheme, annotators are provided with candidate relation instances from distant supervision, and they then manually supplement and remove relational facts based on the recommendations. However, when comparing DocRED with a subset relabeled from scratch, we find that this scheme results in a considerable amount of false negative samples and an obvious bias towards popular entities and relations. Furthermore, we observe that the models trained on DocRED have low recall on our relabeled dataset and inherit the same bias in the training data. Through the analysis of annotators' behaviors, we figure out the underlying reason for the problems above: the scheme actually discourages annotators from supplementing adequate instances in the revision phase. We appeal to future research to take into consideration the issues with the recommend-revise scheme when designing new models and annotation schemes. The relabeled dataset is released at https://github.com/AndrewZhe/ Revisit-DocRED , to serve as a more reliable test set of document RE models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Revisiting DocRED - Addressing the False Negative Problem in Relation ExtractionQingyu Tan, Lu Xu, Lidong Bing, Hwee Tou Ng 等EMNLP 2022 · 被引用 76 次
- A Unified Positive-Unlabeled Learning Framework for Document-Level Relation Extraction with Different Levels of LabelingYe Wang, Xinxin Liu, Wenxin Hu, Tao ZhangEMNLP 2022 · 被引用 18 次
- Anaphor Assisted Document-Level Relation ExtractionChonggang Lu, Richong Zhang, Kai Sun, Jaein Kim 等EMNLP 2023 · 被引用 16 次
- A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete LabelingYe Wang, Huazheng Pan, Tao Zhang, Wen Wu 等AAAI 2024 · 被引用 11 次
- Boosting Document-Level Relation Extraction by Mining and Injecting Logical RulesShengda Fan, Shasha Mo, Jianwei NiuEMNLP 2022 · 被引用 9 次
它引用的顶会 Paper7
- Document-Level Relation Extraction with Adaptive Thresholding and Localized Context PoolingWenxuan Zhou, Kevin Huang, Tengyu Ma, Jing HuangAAAI 2021 · 被引用 360 次
- Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionGuoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei LuACL 2020 · 被引用 294 次
- Double Graph Based Reasoning for Document-level Relation ExtractionShuang Zeng, Runxin Xu, Baobao Chang, Lei LiEMNLP 2020 · 被引用 238 次
- Entity Structure Within and Throughout: Modeling Mention Dependencies for Document-Level Relation ExtractionBenfeng Xu, Quan Wang, Yajuan Lyu, Yong Zhu 等AAAI 2021 · 被引用 200 次
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 被引用 146 次
相关 Paper
- TTM-RE: Memory-Augmented Document-Level Relation ExtractionChufan Gao, Xuan Wang, Jimeng SunACL 2024 · 被引用 11 次
- TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction TaskChristoph Alt, Aleksandra Gabryszak, Leonhard HennigACL 2020 · 被引用 9 次
- SENT: Sentence-level Distant Relation Extraction via Negative TrainingRuotian Ma, Tao Gui, Linyang Li, Qi Zhang 等ACL 2021
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li 等EMNLP 2021 · 被引用 18 次
- Revisiting the Negative Data of Distantly Supervised Relation ExtractionChenhao Xie, Jiaqing Liang, Jingping Liu, Chengsong Huang 等ACL 2021
