Typing Errors in Factual Knowledge Graphs: Severity and Possible Ways Out
Peiran Yao, Denilson Barbosa
Abstract
Large-scale factual knowledge graphs (KGs) such as DBpedia and Wikidata are essential to many popular downstream tasks and are also widely used by various research communities as training and/or benchmarking data. Despite their immense success and utility, these KGs are surprisingly noisy. In this study, we investigate the quality of these KGs, where the typing error rate is estimated to be 27% for coarse-grained types on average, and even 73% for certain fine-grained types. In pursuit of solutions, we propose an active typing error detection algorithm that maximizes the utilization of both gold and noisy labels. We also comprehensively discuss and compare the state-of-the-art in unsupervised, semi-supervised, and supervised paradigms to deal with typing errors in factual KGs. The outcomes of this study provide guidelines for researchers to use noisy factual KGs. To help practitioners deploy the techniques and conduct further research, we published our code and data 1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- An End-To-End Re-Evaluation of Table Entity-LinkersMartin Pekár Christensen, Matteo Lissandrini, Katja HoseICDE 2026
- Knowledge Graph Error Detection with Contrastive Confidence AdaptionXiangyu Liu, Yang Liu, Wei HuAAAI 2024 · 16 citations
- Extraction of Validating Shapes from very large Knowledge GraphsKashif Rabbani, Matteo Lissandrini, Katja HoseVLDB 2023 · 48 citations
- REA: Robust Cross-lingual Entity Alignment Between Knowledge GraphsShichao Pei, Lu Yu, Guoxian Yu, Xiangliang ZhangKDD 2020 · 44 citations
- PGE: Robust Product Graph Embedding Learning for Error DetectionKewei Cheng, Xian Li, Yifan Ethan Xu, Xin Luna Dong et al.VLDB 2022 · 13 citations
