SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text
Miaobo Hu, Xiaobo Guo, Shuhao Hu, BoKun Wang, Rui Chen, Xin Wang, Jun Xiao, Daren Zha
摘要
Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available. We introduce SCOPE (Schema Construction and Ontology-induction Pipeline Evaluation), a train-text-only benchmark for corpus-to-schema induction and optional schema fusion from raw text, built from 24 public information extraction sources (15 RE and 9 EE) normalized into evaluation-only gold schema graphs; its core event-extraction target covers event types and within-event argument roles, with inter-event links reported separately. We present SCION (Schema Construction and Induction with Ontology Normalization), an auditable reference pipeline rather than a new extraction architecture; it constructs candidate spaces from train text and restricts naming, merging, filtering, validation, and conservative fusion to candidate-linked evidence under strict JSON contracts. On the SCOPE core suite, SCION-lite attains the highest F1 among released source-schema references, Text2Onto-style, LLM-only, and matched extract-then-aggregate baselines under Literal, Fuzzy, Continuous, and Graph schema-graph metrics, while the compact open-model SCION-RL variant reduces reliance on proprietary LLM schema engineers. These results are reported against normalized typed-edge targets rather than as claims that induced schemas surpass human ontology design; the release includes evidence-linked outputs, parse/fallback logs, candidate retention/merging logs, run manifests, code, and benchmark packages at https://github.com/wandugu/paper_scion .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- CASIE: Extracting Cybersecurity Event Information from TextTaneeya Satyapanich, Francis Ferraro, Tim FininAAAI 2020 · 被引用 148 次
- What the Role is vs. What Plays the Role: Semi-Supervised Event Argument Extraction via Dual Question AnsweringYang Zhou, Yubo Chen, Jun Zhao, Yin Wu 等AAAI 2021 · 被引用 73 次
- PHEE: A Dataset for Pharmacovigilance Event Extraction from TextZhaoyue Sun, Jiazheng Li, Gabriele Pergola, Byron C. Wallace 等EMNLP 2022 · 被引用 17 次
- SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsQi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea 等EMNLP 2024 · 被引用 7 次
- Multi-Sentence Argument LinkingSeth Ebner, Patrick Xia, Ryan Culkin, Kyle Rawlins 等ACL 2020 · 被引用 1 次
相关 Paper
- Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph ConstructionBowen Zhang, Harold SohEMNLP 2024 · 被引用 65 次
- Linking Surface Facts to Large-Scale Knowledge GraphsGorjan Radevski, Kiril Gashteovski, Chia-Chien Hung, Carolin Lawrence 等EMNLP 2023 · 被引用 2 次
- The Future is not One-dimensional: Complex Event Schema Induction by Graph Modeling for Event PredictionManling Li, Sha Li, Zhenhailong Wang, Lifu Huang 等EMNLP 2021 · 被引用 29 次
- AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale CorporaJiaxin Bai, Wei Fan, Qi Hu, Qing Zong 等ACL 2026 · 被引用 27 次
- SciEvent: Benchmarking Multi-domain Scientific Event ExtractionBofu Dong, Pritesh Shah, Sumedh Sonawane, Tiyasha Banerjee 等EMNLP 2025
