SciREX: A Challenge Dataset for Document-Level Information Extraction
Sarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, Iz Beltagy
摘要
Extracting information from full documents is an important problem in many domains, but most previous work focus on identifying relationships within a sentence or a paragraph. It is challenging to create a large-scale information extraction (IE) dataset at the document level since it requires an understanding of the whole document to annotate entities and their document-level relationships that usually span beyond sentences or even sections. In this paper, we introduce SCIREX, a document level IE dataset that encompasses multiple IE tasks, including salient entity identification and document level N -ary relation identification from scientific articles. We annotate our dataset by integrating automatic and human annotations, leveraging existing scientific knowledge resources. We develop a neural model as a strong baseline that extends previous state-of-the-art IE models to documentlevel IE. Analyzing the model performance shows a significant gap between human performance and current baselines, inviting the community to use our dataset as a challenge to develop document-level IE models. Our data and code are publicly available at https: //github.com/allenai/SciREX
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen 等ICLR 2024 · 被引用 118 次
- MS2: Multi-Document Summarization of Medical StudiesJay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl 等EMNLP 2021 · 被引用 83 次
- End-to-End Argumentation Knowledge Graph ConstructionKhalid Al Khatib, Yufang Hou, Henning Wachsmuth, Charles Jochim 等AAAI 2020 · 被引用 56 次
- Text2NKG: Fine-Grained N-ary Relation Extraction for N-ary relational Knowledge Graph ConstructionHaoran Luo, Haihong E, Yuhao Yang, Tianyu Yao 等NeurIPS 2024 · 被引用 19 次
相关 Paper
- SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsQi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea 等EMNLP 2024 · 被引用 7 次
- SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLPDecheng Duan, Jitong Peng, Yingyi Zhang, Chengzhi ZhangEMNLP 2025 · 被引用 1 次
- Document-level Entity-based Extraction as Template GenerationKung-Hsiang Huang, Sam Tang, Nanyun PengEMNLP 2021 · 被引用 44 次
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li 等EMNLP 2021 · 被引用 18 次
- Global-to-Local Neural Networks for Document-Level Relation ExtractionDifeng Wang, Wei Hu, Ermei Cao, Weijian SunEMNLP 2020 · 被引用 122 次
