Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models
Piotr Przybyla, Euan McGill, Horacio Saggion
Abstract
Large language models have many beneficial applications, but can they also be used to attack content-filtering algorithms in social media platforms? We investigate the challenge of generating adversarial examples to test the robustness of text classification algorithms detecting low-credibility content, including propaganda, false claims, rumours and hyperpartisan news. We focus on simulation of content moderation by setting realistic limits on the number of queries an attacker is allowed to attempt. Within our solution (TREPAT), initial rephrasings are generated by large language models with prompts inspired by meaning-preserving NLP tasks, such as text simplification and style transfer. Subsequently, these modifications are decomposed into small changes, applied through beam search procedure, until the victim classifier changes its decision. We perform (1) quantitative evaluation using various prompts, models and query limits, (2) targeted manual assessment of the generated text and (3) qualitative linguistic analysis. The results confirm the superiority of our approach in the constrained scenario, especially in case of long input text (news articles), where exhaustive search is not feasible.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2bad37f-3900-4355-b47a-da06516d515cCited by top-tier papers2
- Mitigating GenAI-Powered Evidence Pollution for Out-Of-Context Misinformation DetectionZehong Yan, Peng Qi, Wynne Hsu, Mong-Li LeeICDE 2026 · 1 citation
- Adversarial Attacks Against Automated Fact-Checking: A SurveyFanzhen Liu, Sharif Abuadbba, Kristen Moore, Surya Nepal et al.EMNLP 2025
Builds on9
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui et al.ICLR 2024 · 146 citations
- OLMo: Accelerating the Science of Language ModelsDirk Groeneveld, Iz Beltagy, Evan Pete Walsh, Akshita Bhagia et al.ACL 2024 · 52 citations
- Expertise Style Transfer: A New Task Towards Better Communication between Experts and LaymenYixin Cao, Ruihao Shui, Liangming Pan, Min-Yen Kan et al.ACL 2020 · 50 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- RAFT: Realistic Attacks to Fool Text DetectorsJames Wang, Ran Li, Junfeng Yang, Chengzhi MaoEMNLP 2024 · 2 citations
- Large Language Model (LLM)-driven Adversarial Social Influences in Online Information Spread: Risks and InterventionsZhuoran Lu, Gionnieve Lim, Ming YinCHI 2026
- LLM-Generated Fake News Induces Truth Decay in News Ecosystem: A Case Study on Neural News RecommendationBeizhe Hu, Qiang Sheng, Juan Cao, Yang Li et al.SIGIR 2025 · 7 citations
- AI 'News' Content Farms Are Easy to Make and Hard to Detect: A Case Study in ItalianGiovanni Puccetti, Anna Rogers, Chiara Alzetta, Felice Dell'Orletta et al.ACL 2024 · 1 citation
- Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in LegislationAtharvan Dogra, Krishna Pillutla, Ameet Deshpande, Ananya B. Sai et al.ACL 2025
