SciREX: A Challenge Dataset for Document-Level Information Extraction
Sarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, Iz Beltagy
Abstract
Extracting information from full documents is an important problem in many domains, but most previous work focus on identifying relationships within a sentence or a paragraph. It is challenging to create a large-scale information extraction (IE) dataset at the document level since it requires an understanding of the whole document to annotate entities and their document-level relationships that usually span beyond sentences or even sections. In this paper, we introduce SCIREX, a document level IE dataset that encompasses multiple IE tasks, including salient entity identification and document level N -ary relation identification from scientific articles. We annotate our dataset by integrating automatic and human annotations, leveraging existing scientific knowledge resources. We develop a neural model as a strong baseline that extends previous state-of-the-art IE models to documentlevel IE. Analyzing the model performance shows a significant gap between human performance and current baselines, inviting the community to use our dataset as a challenge to develop document-level IE models. Our data and code are publicly available at https: //github.com/allenai/SciREX
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d9e8a16-dff6-478d-bb1b-acd753056b8cCited by top-tier papers27
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity RecognitionWenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen et al.ICLR 2024 · 118 citations
- MS2: Multi-Document Summarization of Medical StudiesJay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl et al.EMNLP 2021 · 83 citations
- End-to-End Argumentation Knowledge Graph ConstructionKhalid Al Khatib, Yufang Hou, Henning Wachsmuth, Charles Jochim et al.AAAI 2020 · 56 citations
- Text2NKG: Fine-Grained N-ary Relation Extraction for N-ary relational Knowledge Graph ConstructionHaoran Luo, Haihong E, Yuhao Yang, Tianyu Yao et al.NeurIPS 2024 · 19 citations
Related papers
- SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsQi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea et al.EMNLP 2024 · 7 citations
- SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLPDecheng Duan, Jitong Peng, Yingyi Zhang, Chengzhi ZhangEMNLP 2025 · 1 citation
- Document-level Entity-based Extraction as Template GenerationKung-Hsiang Huang, Sam Tang, Nanyun PengEMNLP 2021 · 44 citations
- CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildYuan Yao, Jiaju Du, Yankai Lin, Peng Li et al.EMNLP 2021 · 18 citations
- Global-to-Local Neural Networks for Document-Level Relation ExtractionDifeng Wang, Wei Hu, Ermei Cao, Weijian SunEMNLP 2020 · 122 citations
