Unifying Knowledge Base Completion with PU Learning to Mitigate the Observation Bias
Jonas Schouterden, Jessa Bekker, Jesse Davis, Hendrik Blockeel
Abstract
Methods for Knowledge Base Completion (KBC) reason about a knowledge base (KB) in order to derive new facts that should be included in the KB. This is challenging for two reasons. First, KBs only contain positive examples. This complicates model evaluation which needs both positive and negative examples. Second, those facts that were selected to be included in the knowledge base, are most likely not an i.i.d. sample of the true facts, due to the way knowledge bases are constructed. In this paper, we focus on rule-based approaches, which traditionally address the first challenge by making assumptions that enable identifying negative examples, which in turn makes it possible to compute a rule's confidence or precision. However, they largely ignore the second challenge, which means that their estimates of a rule's confidence can be biased. This paper approaches rule-based KBC through the lens of PU-learning, which can cope with both challenges. We make three contributions.: (1) We provide a unifying view that formalizes the relationship between multiple existing confidences measures based on (i) what assumption they make about and (ii) how their accuracy depends on the selection mechanism. (2) We introduce two new confidence measures that can mitigate known biases by using propensity scores that quantify how likely a fact is to be included the KB. (3) We show through theoretical and empirical analysis that taking the bias into account improves the confidence estimates, even when the propensity scores are not known exactly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Recovering the Propensity Score from Biased Positive Unlabeled DataWalter Gerych, Thomas Hartvigsen, Luke Buquicchio, Emmanuel Agu et al.AAAI 2022 · 20 citations
- PUe: Biased Positive-Unlabeled Learning Enhancement by Causal InferenceXutao Wang, Hanting Chen, Tianyu Guo, Yunhe WangNeurIPS 2023 · 10 citations
- Learning from Positive and Unlabeled Data with Arbitrary Positive ShiftZayd Hammoudeh, Daniel LowdNeurIPS 2020 · 53 citations
- Uncertain Knowledge Graph Completion via Semi-Supervised Confidence Distribution LearningTianxing Wu, Shutong Zhu, Jingting Wang, Ning Xu et al.NeurIPS 2025 · 4 citations
- Positive Distribution Pollution: Rethinking Positive Unlabeled Learning from a Unified PerspectiveQianqiao Liang, Mengying Zhu, Yan Wang, Xiuyuan Wang et al.AAAI 2023 · 4 citations
