SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLP
Decheng Duan, Jitong Peng, Yingyi Zhang, Chengzhi Zhang
Abstract
Structured information extraction from scientific literature is crucial for capturing core concepts and emerging trends in specialized fields. While existing datasets aid model development, most focus on specific publication sections due to domain complexity and the high cost of annotating scientific texts. To address this limitation, we introduce SciNLP-a specialized benchmark for full-text entity and relation extraction in the Natural Language Processing (NLP) domain. The dataset comprises 60 manually annotated full-text NLP publications, covering 6,409 entities and 1,648 relations. Compared to existing research, SciNLP is the first dataset providing full-text annotations of entities and their relationships in the NLP domain. To validate the effectiveness of SciNLP, we conducted comparative experiments with similar datasets and evaluated the performance of stateof-the-art supervised models on this dataset. Results reveal varying extraction capabilities of existing models across academic texts of different lengths. Cross-comparisons with existing datasets show that SciNLP achieves significant performance improvements on certain baseline models. Using models trained on SciNLP, we implemented automatic construction of a finegrained knowledge graph for the NLP domain. Our KG has an average node degree of 3.3 per entity, indicating rich semantic topological information that enhances downstream applications. The dataset is publicly available at: https://github.com/AKADDC/SciNLP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0cbd50f2-b7ea-495d-a0d5-154f5d70f50bBuilds on8
- A Partition Filter Network for Joint Entity and Relation ExtractionZhiheng Yan, Chong Zhang, Jinlan Fu, Qi Zhang et al.EMNLP 2021 · 142 citations
- Packed Levitated Marker for Entity and Relation ExtractionDeming Ye, Yankai Lin, Peng Li, Maosong SunACL 2022 · 140 citations
- DEKR: Description Enhanced Knowledge Graph for Machine Learning Method RecommendationXianshuai Cao, Yuliang Shi, Han Yu, Jihu Wang et al.SIGIR 2021 · 33 citations
- SciMON: Scientific Inspiration Machines Optimized for NoveltyQingyun Wang, Doug Downey, Heng Ji, Tom HopeACL 2024 · 22 citations
- PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity RecognitionTao Zhang, Congying Xia, Philip S. Yu, Zhiwei Liu et al.EMNLP 2021 · 22 citations
Related papers
- SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsQi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea et al.EMNLP 2024 · 7 citations
- SciREX: A Challenge Dataset for Document-Level Information ExtractionSarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, Iz BeltagyACL 2020 · 9 citations
- End-to-End Argumentation Knowledge Graph ConstructionKhalid Al Khatib, Yufang Hou, Henning Wachsmuth, Charles Jochim et al.AAAI 2020 · 56 citations
- SciNLI: A Corpus for Natural Language Inference on Scientific TextMobashir Sadat, Cornelia CarageaACL 2022 · 41 citations
- GSAP-ERE: Fine-Grained Scholarly Entity and Relation Extraction Focused on Machine LearningWolfgang Otto, Lu Gan, Sharmila Upadhyaya, Saurav Karmakar et al.AAAI 2026
