SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text
Miaobo Hu, Xiaobo Guo, Shuhao Hu, BoKun Wang, Rui Chen, Xin Wang, Jun Xiao, Daren Zha
Abstract
Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available. We introduce SCOPE (Schema Construction and Ontology-induction Pipeline Evaluation), a train-text-only benchmark for corpus-to-schema induction and optional schema fusion from raw text, built from 24 public information extraction sources (15 RE and 9 EE) normalized into evaluation-only gold schema graphs; its core event-extraction target covers event types and within-event argument roles, with inter-event links reported separately. We present SCION (Schema Construction and Induction with Ontology Normalization), an auditable reference pipeline rather than a new extraction architecture; it constructs candidate spaces from train text and restricts naming, merging, filtering, validation, and conservative fusion to candidate-linked evidence under strict JSON contracts. On the SCOPE core suite, SCION-lite attains the highest F1 among released source-schema references, Text2Onto-style, LLM-only, and matched extract-then-aggregate baselines under Literal, Fuzzy, Continuous, and Graph schema-graph metrics, while the compact open-model SCION-RL variant reduces reliance on proprietary LLM schema engineers. These results are reported against normalized typed-edge targets rather than as claims that induced schemas surpass human ontology design; the release includes evidence-linked outputs, parse/fallback logs, candidate retention/merging logs, run manifests, code, and benchmark packages at https://github.com/wandugu/paper_scion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 309f24c0-e0e1-4203-9b15-1c4d1c055f76Builds on5
- CASIE: Extracting Cybersecurity Event Information from TextTaneeya Satyapanich, Francis Ferraro, Tim FininAAAI 2020 · 148 citations
- What the Role is vs. What Plays the Role: Semi-Supervised Event Argument Extraction via Dual Question AnsweringYang Zhou, Yubo Chen, Jun Zhao, Yin Wu et al.AAAI 2021 · 73 citations
- PHEE: A Dataset for Pharmacovigilance Event Extraction from TextZhaoyue Sun, Jiazheng Li, Gabriele Pergola, Byron C. Wallace et al.EMNLP 2022 · 17 citations
- SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsQi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea et al.EMNLP 2024 · 7 citations
- Multi-Sentence Argument LinkingSeth Ebner, Patrick Xia, Ryan Culkin, Kyle Rawlins et al.ACL 2020 · 1 citation
Related papers
- Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph ConstructionBowen Zhang, Harold SohEMNLP 2024 · 65 citations
- Linking Surface Facts to Large-Scale Knowledge GraphsGorjan Radevski, Kiril Gashteovski, Chia-Chien Hung, Carolin Lawrence et al.EMNLP 2023 · 2 citations
- The Future is not One-dimensional: Complex Event Schema Induction by Graph Modeling for Event PredictionManling Li, Sha Li, Zhenhailong Wang, Lifu Huang et al.EMNLP 2021 · 29 citations
- AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale CorporaJiaxin Bai, Wei Fan, Qi Hu, Qing Zong et al.ACL 2026 · 27 citations
- SciEvent: Benchmarking Multi-domain Scientific Event ExtractionBofu Dong, Pritesh Shah, Sumedh Sonawane, Tiyasha Banerjee et al.EMNLP 2025
