Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
Haorui He, Yupeng Li, Bin Benjamin Zhu, Dacheng Wen, Reynold Cheng, Francis C. M. Lau
Abstract
State-of-the-art (SOTA) fact-checking systems combat misinformation by employing autonomous LLM-based agents to decompose complex claims into smaller sub-claims, verify each sub-claim individually, and aggregate the partial results to produce verdicts with justifications (explanations for the verdicts). The security of these systems is crucial, as compromised fact-checkers can amplify misinformation, but remains largely underexplored. To bridge this gap, this work introduces a novel threat model against such fact-checking systems and presents Fact2Fiction, the first poisoning attack framework targeting SOTA agentic fact-checking systems. Fact2Fiction employs LLMs to mimic the decomposition strategy and exploit system-generated justifications to craft tailored malicious evidences that compromise sub-claim verification. Extensive experiments demonstrate that Fact2Fiction achieves 8.9%-21.2% higher attack success rates than SOTA attacks across various poisoning budgets and exposes security weaknesses in existing fact-checking systems, highlighting the need for defensive countermeasures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2601ade-a82e-4719-85b8-a44761cf7c84Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- Synthetic Disinformation Attacks on Automated Fact Verification SystemsYibing Du, Antoine Bosselut, Christopher D. ManningAAAI 2022 · 58 citations
- Generating Literal and Implied Subquestions to Fact-check Complex ClaimsJifan Chen, Aniruddh Sriram, Eunsol Choi, Greg DurrettEMNLP 2022 · 30 citations
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman et al.KDD 2025 · 27 citations
- DEFAME: Dynamic Evidence-based FAct-checking with Multimodal ExpertsTobias Braun, Mark Rothermel, Marcus Rohrbach, Anna RohrbachICML 2025
Related papers
- LoCal: Logical and Causal Fact-Checking with LLM-Based Multi-AgentsJiatong Ma, Linmei Hu, Rang Li, Wenbo FuWWW 2025 · 31 citations
- Fact-Saboteurs: A Taxonomy of Evidence Manipulation Attacks against Fact-Verification SystemsSahar Abdelnabi, Mario FritzUSENIX Security 2023
- FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language ModelsHongzhan Lin, Yang Deng, Yuxuan Gu, Wenxuan Zhang et al.ACL 2025
- MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Knowledge Poisoning AttacksHyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim, Dimitrios Bralios et al.ACL 2026 · 1 citation
- Multi-Sourced, Multi-Agent Evidence Retrieval for Fact-CheckingShuzhi Gong, Richard O. Sinnott, Jianzhong Qi, Cécile Paris et al.SIGIR 2026 · 2 citations
