ACL2026

Large Language Models Require Curated Context for Reliable Political Fact-Checking - Even with Reasoning and Web Search

Matthew R. DeVerna, Kai-Cheng Yang, Harry Yaojun Yan, Filippo Menczer

4 citations

Abstract

Large language models (LLMs) have raised hopes for automated end-to-end fact-checking, but prior studies report mixed results. As mainstream chatbots increasingly ship with reasoning capabilities and web search toolsand millions of users already rely on them for verification-rigorous evaluation is urgent. We evaluate 15 recent LLMs from OpenAI, Google, Meta, and DeepSeek on more than 6,000 claims fact-checked by PolitiFact, comparing standard models with reasoning-and web-search variants. Standard models perform poorly, reasoning offers minimal benefits, and web search provides only moderate gains, despite fact-checks being available on the web. In contrast, a curated RAG system using PolitiFact summaries improved macro F1 by 233% on average across model variants. These findings suggest that giving models access to curated high-quality context is a promising path for automated factchecking. 2017), RumorEval (Gorrell et al., 2019), CLEF CheckThat! (Nakov et al., 2022), and Claim-Buster (Hassan et al., 2017) spurred this line of work. Transformer-based models such as BERT produced major accuracy gains on these benchmarks (Soleimani et al., 2020; Nie et al., 2019; Zeng et al., 2021) . The emergence of LLMs has recently opened the door to effective end-to-end fact-checking. We examine political fact-checking: how well AI systems reproduce the veracity labels assigned by professional fact-checkers. Hoes et al. (2023) reported nearly 70% accuracy for GPT-3 on PolitiFact claims when labels were collapsed to