Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond
Georgios Karagiannis, Immanuel Trummer, Saehan Jo, Shubham Khandelwal, Xuezhi Wang, Cong Yu
摘要
We introduce the problem of anti-knowledge mining. Our goal is to create an "anti-knowledge base" that contains factual mistakes. The resulting data can be used for analysis, training, and benchmarking in the research domain of automated fact checking. Prior data sets feature manually generated fact checks of famous misclaims. Instead, we focus on the long tail of factual mistakes made by Web authors, ranging from erroneous sports results to incorrect capitals.
We mine mistakes automatically, by an unsupervised approach, from Wikipedia updates that correct factual mistakes. Identifying such updates (only a small fraction of the total number of updates) is one of the primary challenges. We mine anti-knowledge by a multi-step pipeline. First, we filter out candidate updates via several simple heuristics. Next, we correlate Wikipedia updates with other statements made on the Web. Using claim occurrence frequencies as input to a probabilistic model, we infer the likelihood of corrections via an iterative expectation-maximization approach. Finally, we extract mistakes in the form of subjectpredicate-object triples and rank them according to several criteria. Our end result is a data set containing over 110,000 ranked mistakes with a precision of 85% in the top 1% and a precision of over 60% in the top 25%. We demonstrate that baselines achieve significantly lower precision. Also, we exploit our data to verify several hypothesis on why users make mistakes. We finally show that the AKB can be used to find mistakes on the entire Web.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DETERRENT: Knowledge Guided Graph Attention Network for Detecting Healthcare MisinformationLimeng Cui, Haeseung Seo, Maryam Tabar, Fenglong Ma 等KDD 2020 · 被引用 163 次
- Longitudinal Assessment of Reference Quality on WikipediaAitolkyn Baigutanova, Jaehyeon Myung, Diego Sáez-Trumper, Ai-Jou Chou 等WWW 2023 · 被引用 14 次
- CEDAR: A System for Cost-Efficient Data-Driven Claim VerificationTharushi Jayasekara, Immanuel TrummerVLDB 2025 · 被引用 1 次
- Scrutinizer: A Mixed-Initiative Approach to Large-Scale, Data-Driven Claim VerificationGeorgios Karagiannis, Mohammed Saeed, Paolo Papotti, Immanuel TrummerVLDB 2020
相关 Paper
- AKEW: Assessing Knowledge Editing in the WildXiaobao Wu, Liangming Pan, William Yang Wang, Anh Tuan LuuEMNLP 2024 · 被引用 2 次
- Open Knowledge Enrichment for Long-tail EntitiesErmei Cao, Difeng Wang, Jiacheng Huang, Wei HuWWW 2020 · 被引用 51 次
- Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language ModelsSina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua, Thanawan Atchariyachanvanit 等EMNLP 2025
- Correcting Knowledge Base AssertionsJiaoyan Chen, Xi Chen, Ian Horrocks, Erik B. Myklebust 等WWW 2020 · 被引用 23 次
- Typing Errors in Factual Knowledge Graphs: Severity and Possible Ways OutPeiran Yao, Denilson BarbosaWWW 2021 · 被引用 8 次
