SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific Documents
Qi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea, Longin Jan Latecki, Eduard C. Dragut
Abstract
Scientific information extraction (SciIE) is critical for converting unstructured knowledge from scholarly articles into structured data (entities and relations). Several datasets have been proposed for training and validating SciIE models. However, due to the high complexity and cost of annotating scientific texts, those datasets restrict their annotations to specific parts of paper, such as abstracts, resulting in the loss of diverse entity mentions and relations in context. In this paper, we release a new entity and relation extraction dataset for entities related to datasets, methods, and tasks in scientific articles. Our dataset contains 106 manually annotated full-text scientific publications with over 24k entities and 12k relations. To capture the intricate use and interactions among entities in full texts, our dataset contains a finegrained tag set for relations. Additionally, we provide an out-of-distribution test set to offer a more realistic evaluation. We conduct comprehensive experiments, including state-of-the-art supervised models and our proposed LLM baselines, and highlight the challenges presented by our dataset, encouraging the development of innovative models to further the field of SciIE.1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLPDecheng Duan, Jitong Peng, Yingyi Zhang, Chengzhi ZhangEMNLP 2025 · 1 citation
- HGNet: Scalable Foundation Model for Automated Knowledge Graph Generation from Scientific LiteratureDevvrat Joshi, Islem RekikICLR 2026 · 1 citation
- SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from TextMiaobo Hu, Xiaobo Guo, Shuhao Hu, BoKun Wang et al.ICML 2026
- GSAP-ERE: Fine-Grained Scholarly Entity and Relation Extraction Focused on Machine LearningWolfgang Otto, Lu Gan, Sharmila Upadhyaya, Saurav Karmakar et al.AAAI 2026
- SciEvent: Benchmarking Multi-domain Scientific Event ExtractionBofu Dong, Pritesh Shah, Sumedh Sonawane, Tiyasha Banerjee et al.EMNLP 2025
Builds on15
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney et al.ACL 2020 · 424 citations
- GoLLIE: Annotation Guidelines improve Zero-Shot Information-ExtractionOscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle et al.ICLR 2024 · 168 citations
- Revisiting Relation Extraction in the era of Large Language ModelsSomin Wadhwa, Silvio Amir, Byron C. WallaceACL 2023 · 145 citations
Related papers
- SciREX: A Challenge Dataset for Document-Level Information ExtractionSarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, Iz BeltagyACL 2020 · 9 citations
- WebIE: Faithful and Robust Information Extraction on the WebChenxi Whitehouse, Clara Vania, Alham Fikri Aji, Christos Christodoulopoulos et al.ACL 2023 · 3 citations
- REDFM: a Filtered and Multilingual Relation Extraction DatasetPere-Lluís Huguet Cabot, Simone Tedeschi, Axel-Cyrille Ngonga Ngomo, Roberto NavigliACL 2023 · 9 citations
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationLiang He, Hongke Wang, Yongchang Cao, Zhen Wu et al.ACM MM 2023 · 17 citations
- Modeling Entities as Semantic Points for Visual Information Extraction in the WildZhibo Yang, Rujiao Long, Pengfei Wang, Sibo Song et al.CVPR 2023
