What is Normal, What is Strange, and What is Missing in a Knowledge Graph: Unified Characterization via Inductive Summarization
Caleb Belth, Xinyi Zheng, Jilles Vreeken, Danai Koutra
Abstract
Knowledge graphs (KGs) store highly heterogeneous information about the world in the structure of a graph, and are useful for tasks such as question answering and reasoning. However, they often contain errors and are missing information. Vibrant research in KG refinement has worked to resolve these issues, tailoring techniques to either detect specific types of errors or complete a KG. In this work, we introduce a unified solution to KG characterization by formulating the problem as unsupervised KG summarization with a set of inductive, soft rules, which describe what is normal in a KG, and thus can be used to identify what is abnormal, whether it be strange or missing. Unlike first-order logic rules, our rules are labeled, rooted graphs, i.e., patterns that describe the expected neighborhood around a (seen or unseen) node, based on its type, and information in the KG. Stepping away from the traditional support/confidence-based rule mining techniques, we propose KGist, Knowledge Graph Inductive SummarizaTion, which learns a summary of inductive rules that best compress the KG according to the Minimum Description Length principle—a formulation that we are the first to use in the context of KG rule mining. We apply our rules to three large KGs (NELL, DBpedia, and Yago), and tasks such as compression, various types of error detection, and identification of incomplete information. We show that KGist outperforms task-specific, supervised and unsupervised baselines in error detection and incompleteness identification, (identifying the location of up to 93% of missing entities—over 10% more than baselines), while also being efficient for large knowledge graphs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d05c09e1-dc57-4fd2-8270-d110ababb8c1Cited by top-tier papers6
- MulDE: Multi-teacher Knowledge Distillation for Low-dimensional Knowledge Graph EmbeddingsKai Wang, Yu Liu, Qian Ma, Quan Z. ShengWWW 2021 · 67 citations
- Multimodal Knowledge Graph Error Detection with Disentanglement VAE and Multi-Grained Triplet ConfidenceXuhui Sui, Ying Zhang, Yu Zhao, Baohang Zhou et al.WWW 2025 · 2 citations
- Online Detection of Anomalies in Temporal Knowledge Graphs with InterpretabilityJiasheng Zhang, Rex Ying, Jie ShaoSIGMOD 2025 · 2 citations
- APEX2: Adaptive and Extreme Summarization for Personalized Knowledge GraphsZihao Li, Dongqi Fu, Mengting Ai, Jingrui HeKDD 2025 · 1 citation
- Evaluating the Calibration of Knowledge Graph Embeddings for Trustworthy Link PredictionTara Safavi, Danai Koutra, Edgar MeijEMNLP 2020 · 1 citation
Related papers
- Towards Global-Topology Relation Graph for Inductive Knowledge Graph CompletionLing Ding, Lei Huang, Zhizhi Yu, Di Jin et al.AAAI 2025 · 8 citations
- AdaRPT: An Adaptive Rule Pattern Transfer Model for Fully Inductive Knowledge Graph ReasoningZhiwen Xie, Zhuo Zhao, Jinjin Ma, Guangyou Zhou et al.SIGIR 2025 · 3 citations
- Generating Graph-Like Logical Rules for Knowledge Graph Reasoning via Diffusion ModelsHaoxiang Cheng, Yunfei Wang, Chao Chen, Kewei Cheng et al.KDD 2026 · 1 citation
- Inductive Logical Query Answering in Knowledge GraphsMichael Galkin, Zhaocheng Zhu, Hongyu Ren, Jian TangNeurIPS 2022 · 36 citations
- InGram: Inductive Knowledge Graph Embedding via Relation GraphsJaejun Lee, Chanyoung Chung, Joyce Jiyoung WhangICML 2023 · 83 citations
