Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond
Georgios Karagiannis, Immanuel Trummer, Saehan Jo, Shubham Khandelwal, Xuezhi Wang, Cong Yu
Abstract
We introduce the problem of anti-knowledge mining. Our goal is to create an "anti-knowledge base" that contains factual mistakes. The resulting data can be used for analysis, training, and benchmarking in the research domain of automated fact checking. Prior data sets feature manually generated fact checks of famous misclaims. Instead, we focus on the long tail of factual mistakes made by Web authors, ranging from erroneous sports results to incorrect capitals.
We mine mistakes automatically, by an unsupervised approach, from Wikipedia updates that correct factual mistakes. Identifying such updates (only a small fraction of the total number of updates) is one of the primary challenges. We mine anti-knowledge by a multi-step pipeline. First, we filter out candidate updates via several simple heuristics. Next, we correlate Wikipedia updates with other statements made on the Web. Using claim occurrence frequencies as input to a probabilistic model, we infer the likelihood of corrections via an iterative expectation-maximization approach. Finally, we extract mistakes in the form of subjectpredicate-object triples and rank them according to several criteria. Our end result is a data set containing over 110,000 ranked mistakes with a precision of 85% in the top 1% and a precision of over 60% in the top 25%. We demonstrate that baselines achieve significantly lower precision. Also, we exploit our data to verify several hypothesis on why users make mistakes. We finally show that the AKB can be used to find mistakes on the entire Web.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08333209-0274-429b-946e-01b403cf1fc2Cited by top-tier papers4
- DETERRENT: Knowledge Guided Graph Attention Network for Detecting Healthcare MisinformationLimeng Cui, Haeseung Seo, Maryam Tabar, Fenglong Ma et al.KDD 2020 · 163 citations
- Longitudinal Assessment of Reference Quality on WikipediaAitolkyn Baigutanova, Jaehyeon Myung, Diego Sáez-Trumper, Ai-Jou Chou et al.WWW 2023 · 14 citations
- CEDAR: A System for Cost-Efficient Data-Driven Claim VerificationTharushi Jayasekara, Immanuel TrummerVLDB 2025 · 1 citation
- Scrutinizer: A Mixed-Initiative Approach to Large-Scale, Data-Driven Claim VerificationGeorgios Karagiannis, Mohammed Saeed, Paolo Papotti, Immanuel TrummerVLDB 2020
Related papers
- AKEW: Assessing Knowledge Editing in the WildXiaobao Wu, Liangming Pan, William Yang Wang, Anh Tuan LuuEMNLP 2024 · 2 citations
- Open Knowledge Enrichment for Long-tail EntitiesErmei Cao, Difeng Wang, Jiacheng Huang, Wei HuWWW 2020 · 51 citations
- Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language ModelsSina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua, Thanawan Atchariyachanvanit et al.EMNLP 2025
- Correcting Knowledge Base AssertionsJiaoyan Chen, Xi Chen, Ian Horrocks, Erik B. Myklebust et al.WWW 2020 · 23 citations
- Typing Errors in Factual Knowledge Graphs: Severity and Possible Ways OutPeiran Yao, Denilson BarbosaWWW 2021 · 8 citations
