Can We Predict New Facts with Open Knowledge Graph Embeddings? A Benchmark for Open Link Prediction
Samuel Broscheit, Kiril Gashteovski, Yanjie Wang, Rainer Gemulla
Abstract
Open Information Extraction systems extract ("subject text", "relation text", "object text") triples from raw text. Some triples are textual versions of facts, i.e., non-canonicalized mentions of entities and relations. In this paper, we investigate whether it is possible to infer new facts directly from the open knowledge graph without any canonicalization or any supervision from curated knowledge. For this purpose, we propose the open link prediction task, i.e., predicting test facts by completing ("subject text", "relation text", ?) questions. An evaluation in such a setup raises the question if a correct prediction is actually a new fact that was induced by reasoning over the open knowledge graph or if it can be trivially explained. For example, facts can appear in different paraphrased textual variants, which can lead to test leakage. To this end, we propose an evaluation protocol and a methodology for creating the open link prediction benchmark OLPBENCH. We performed experiments with a prototypical knowledge graph embedding model for open link prediction. While the task is very challenging, our results suggests that it is possible to predict genuinely new facts, which can not be trivially explained.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e26b6550-5b0e-4d7f-a19e-2ea65af33205Cited by top-tier papers6
- Disentangling Sampling and Labeling Bias for Learning in Large-output SpacesAnkit Singh Rawat, Aditya Krishna Menon, Wittawat Jitkrittum, Sadeep Jayasumana et al.ICML 2021 · 13 citations
- Linking Surface Facts to Large-Scale Knowledge GraphsGorjan Radevski, Kiril Gashteovski, Chia-Chien Hung, Carolin Lawrence et al.EMNLP 2023 · 2 citations
- On Synthesizing Data for Context Attribution in Question AnsweringGorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon et al.ACL 2025 · 1 citation
- "Covid vaccine is against Covid but Oxford vaccine is made at Oxford!" Semantic Interpretation of Proper Noun CompoundsKeshav Kolluru, Gabriel Stanovsky, MausamEMNLP 2022 · 1 citation
- BenchIE: A Framework for Multi-Faceted Fact-Based Open Information Extraction EvaluationKiril Gashteovski, Mingying Yu, Bhushan Kotnis, Carolin Lawrence et al.ACL 2022
Related papers
- Joint Open Knowledge Base Canonicalization and LinkingYinan Liu, Wei Shen, Yuanfei Wang, Jianyong Wang et al.SIGMOD 2021 · 16 citations
- Knowledge Base Completion Meets Transfer LearningVid Kocijan, Thomas LukasiewiczEMNLP 2021
- Realistic Re-evaluation of Knowledge Graph Completion Methods: An Experimental StudyFarahnaz Akrami, Mohammed Samiul Saeef, Qingheng Zhang, Wei Hu et al.SIGMOD 2020 · 101 citations
- IELM: An Open Information Extraction Benchmark for Pre-Trained Language ModelsChenguang Wang, Xiao Liu, Dawn SongEMNLP 2022 · 3 citations
- CoDEx: A Comprehensive Knowledge Graph Completion BenchmarkTara Safavi, Danai KoutraEMNLP 2020 · 97 citations
