Fact-Saboteurs: A Taxonomy of Evidence Manipulation Attacks against Fact-Verification Systems
Sahar Abdelnabi, Mario Fritz
摘要
Mis- and disinformation are a substantial global threat to our security and safety. To cope with the scale of online misinformation, researchers have been working on automating fact-checking by retrieving and verifying against relevant evidence. However, despite many advances, a comprehensive evaluation of the possible attack vectors against such systems is still lacking. Particularly, the automated fact-verification process might be vulnerable to the exact disinformation campaigns it is trying to combat. In this work, we assume an adversary that automatically tampers with the online evidence in order to disrupt the fact-checking model via camouflaging the relevant evidence or planting a misleading one. We first propose an exploratory taxonomy that spans these two targets and the different threat model dimensions. Guided by this, we design and propose several potential attack methods. We show that it is possible to subtly modify claim-salient snippets in the evidence and generate diverse and claim-aligned evidence. Thus, we highly degrade the fact-checking performance under many different permutations of the taxonomy's dimensions. The attacks are also robust against post-hoc modifications of the claim. Our analysis further hints at potential limitations in models' inference when faced with contradicting evidence. We emphasize that these attacks can have harmful implications on the inspectable and human-in-the-loop usage scenarios of such models, and we conclude by discussing challenges and directions for future defenses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Mitigating GenAI-Powered Evidence Pollution for Out-Of-Context Misinformation DetectionZehong Yan, Peng Qi, Wynne Hsu, Mong-Li LeeICDE 2026 · 被引用 1 次
- Evaluating LLM-based Personal Information Extraction and CountermeasuresYupei Liu, Yuqi Jia, Jinyuan Jia, Neil Zhenqiang GongUSENIX Security 2025
- What Evidence Do Language Models Find Convincing?Alexander Wan, Eric Wallace, Dan KleinACL 2024
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- Factuality Enhanced Language Models for Open-Ended Text GenerationNayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary 等NeurIPS 2022 · 被引用 318 次
- Coreferential Reasoning Learning for Language RepresentationDeming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu 等EMNLP 2020 · 被引用 164 次
相关 Paper
- Synthetic Disinformation Attacks on Automated Fact Verification SystemsYibing Du, Antoine Bosselut, Christopher D. ManningAAAI 2022 · 被引用 58 次
- Adversarial Attacks Against Automated Fact-Checking: A SurveyFanzhen Liu, Sharif Abuadbba, Kristen Moore, Surya Nepal 等EMNLP 2025
- Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking SystemHaorui He, Yupeng Li, Bin Benjamin Zhu, Dacheng Wen 等AAAI 2026 · 被引用 3 次
- Factoring Fact-Checks: Structured Information Extraction from Fact-Checking ArticlesShan Jiang, Simon Baumgartner, Abe Ittycheriah, Cong YuWWW 2020 · 被引用 28 次
- Automated Justification Production for Claim Veracity in Fact Checking: A Survey on Architectures and ApproachesIslam Eldifrawi, Shengrui Wang, Amine TrabelsiACL 2024 · 被引用 4 次
