Revisiting DocRED - Addressing the False Negative Problem in Relation Extraction
Qingyu Tan, Lu Xu, Lidong Bing, Hwee Tou Ng, Sharifah Mahani Aljunied
摘要
The DocRED dataset is one of the most popular and widely used benchmarks for documentlevel relation extraction (RE). It adopts a recommend-revise annotation scheme so as to have a large-scale annotated dataset. However, we find that the annotation of DocRED is incomplete, i.e., false negative samples are prevalent. We analyze the causes and effects of the overwhelming false negative problem in the DocRED dataset. To address the shortcoming, we re-annotate 4,053 documents in the DocRED dataset by adding the missed relation triples back to the original DocRED. We name our revised DocRED dataset Re-DocRED. We conduct extensive experiments with state-ofthe-art neural models on both datasets, and the experimental results show that the models trained and evaluated on our Re-DocRED achieve performance improvements of around 13 F1 points. Moreover, we conduct a comprehensive analysis to identify the potential areas for further improvement. 1 * Equal contribution. Qingyu Tan and Lu Xu are under the Joint PhD Program between Alibaba and NUS/SUTD. † Corresponding author. 1 Our dataset is publicly available at https://github. com/tonytan48/Re-DocRED .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph ConstructionBowen Zhang, Harold SohEMNLP 2024 · 被引用 65 次
- Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet ExtractionQi Sun, Kun Huang, Xiaocui Yang, Rong Tong 等WWW 2024 · 被引用 40 次
- A Novel Table-to-Graph Generation Approach for Document-Level Joint Entity and Relation ExtractionRuoyu Zhang, Yanzeng Li, Lei ZouACL 2023 · 被引用 20 次
- Revisiting Document-Level Relation Extraction with Context-Guided Link PredictionMonika Jain, Raghava Mutharaju, Ramakanth Kavuluru, Kuldeep SinghAAAI 2024 · 被引用 17 次
- A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete LabelingYe Wang, Huazheng Pan, Tao Zhang, Wen Wu 等AAAI 2024 · 被引用 11 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Document-Level Relation Extraction with Adaptive Thresholding and Localized Context PoolingWenxuan Zhou, Kevin Huang, Tengyu Ma, Jing HuangAAAI 2021 · 被引用 360 次
- Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionGuoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei LuACL 2020 · 被引用 294 次
- Effective Modeling of Encoder-Decoder Architecture for Joint Entity and Relation ExtractionTapas Nayak, Hwee Tou NgAAAI 2020 · 被引用 272 次
- MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERRan Zhou, Xin Li, Ruidan He, Lidong Bing 等ACL 2022 · 被引用 114 次
相关 Paper
- Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocREDQuzhe Huang, Shibo Hao, Yuan Ye, Shengqi Zhu 等ACL 2022
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 被引用 146 次
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li 等EMNLP 2021 · 被引用 18 次
- TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction TaskChristoph Alt, Aleksandra Gabryszak, Leonhard HennigACL 2020 · 被引用 9 次
- TTM-RE: Memory-Augmented Document-Level Relation ExtractionChufan Gao, Xuan Wang, Jimeng SunACL 2024 · 被引用 11 次
