Revisiting DocRED - Addressing the False Negative Problem in Relation Extraction
Qingyu Tan, Lu Xu, Lidong Bing, Hwee Tou Ng, Sharifah Mahani Aljunied
Abstract
The DocRED dataset is one of the most popular and widely used benchmarks for documentlevel relation extraction (RE). It adopts a recommend-revise annotation scheme so as to have a large-scale annotated dataset. However, we find that the annotation of DocRED is incomplete, i.e., false negative samples are prevalent. We analyze the causes and effects of the overwhelming false negative problem in the DocRED dataset. To address the shortcoming, we re-annotate 4,053 documents in the DocRED dataset by adding the missed relation triples back to the original DocRED. We name our revised DocRED dataset Re-DocRED. We conduct extensive experiments with state-ofthe-art neural models on both datasets, and the experimental results show that the models trained and evaluated on our Re-DocRED achieve performance improvements of around 13 F1 points. Moreover, we conduct a comprehensive analysis to identify the potential areas for further improvement. 1 * Equal contribution. Qingyu Tan and Lu Xu are under the Joint PhD Program between Alibaba and NUS/SUTD. † Corresponding author. 1 Our dataset is publicly available at https://github. com/tonytan48/Re-DocRED .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99decbda-3072-4932-8990-d5d13158a2d3Cited by top-tier papers17
- Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph ConstructionBowen Zhang, Harold SohEMNLP 2024 · 65 citations
- Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet ExtractionQi Sun, Kun Huang, Xiaocui Yang, Rong Tong et al.WWW 2024 · 40 citations
- A Novel Table-to-Graph Generation Approach for Document-Level Joint Entity and Relation ExtractionRuoyu Zhang, Yanzeng Li, Lei ZouACL 2023 · 20 citations
- Revisiting Document-Level Relation Extraction with Context-Guided Link PredictionMonika Jain, Raghava Mutharaju, Ramakanth Kavuluru, Kuldeep SinghAAAI 2024 · 17 citations
- A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete LabelingYe Wang, Huazheng Pan, Tao Zhang, Wen Wu et al.AAAI 2024 · 11 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Document-Level Relation Extraction with Adaptive Thresholding and Localized Context PoolingWenxuan Zhou, Kevin Huang, Tengyu Ma, Jing HuangAAAI 2021 · 360 citations
- Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionGuoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei LuACL 2020 · 294 citations
- Effective Modeling of Encoder-Decoder Architecture for Joint Entity and Relation ExtractionTapas Nayak, Hwee Tou NgAAAI 2020 · 272 citations
- MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERRan Zhou, Xin Li, Ruidan He, Lidong Bing et al.ACL 2022 · 114 citations
Related papers
- Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocREDQuzhe Huang, Shibo Hao, Yuan Ye, Shengqi Zhu et al.ACL 2022
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 146 citations
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li et al.EMNLP 2021 · 18 citations
- TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction TaskChristoph Alt, Aleksandra Gabryszak, Leonhard HennigACL 2020 · 9 citations
- TTM-RE: Memory-Augmented Document-Level Relation ExtractionChufan Gao, Xuan Wang, Jimeng SunACL 2024 · 11 citations
