Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
Sungwon Park, Sungwon Han, Xing Xie, Jae-Gil Lee, Meeyoung Cha
Abstract
The spread of fake news harms individuals and presents a critical social challenge that must be addressed. Although numerous algorithmic and insightful features have been developed to detect fake news, many of these features can be manipulated with styleconversion attacks, especially with the emergence of advanced language models, making it more difficult to differentiate from genuine news. This study proposes adversarial style augmentation, AdStyle, designed to train a fake news detector that remains robust against various style-conversion attacks. The primary mechanism involves the strategic use of LLMs to automatically generate a diverse and coherent array of style-conversion attack prompts, enhancing the generation of particularly challenging prompts for the detector. Experiments indicate that our augmentation strategy significantly improves robustness and detection performance when evaluated on fake news benchmark datasets. CCS Concepts • Security and privacy → Software and application security; • Computing methodologies → Artificial intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Generate First, Then Sample: Enhancing Fake News Detection with LLM-Augmented Reinforced SamplingZhao Tong, Yimeng Gu, Huidong Liu, Qiang Liu et al.ACL 2025 · 14 citations
- LLM-Generated Fake News Induces Truth Decay in News Ecosystem: A Case Study on Neural News RecommendationBeizhe Hu, Qiang Sheng, Juan Cao, Yang Li et al.SIGIR 2025 · 7 citations
- Robust Fake News Detection using Large Language Models under Adversarial Sentiment AttacksSahar Tahmasebi, Eric Müller-Budack, Ralph EwerthWWW 2026 · 3 citations
- SGT: Securing Open-Source LLMs Against Malicious Fine-tuning via Safety Guidance TriggerSunguk Shin, Fangzhao Wu, Byung-Jun Lee, Meeyoung Cha et al.ACL 2026
- A Symbolic Adversarial Learning Framework for Evolving Fake News Generation and DetectionChong Tian, Qirong Ho, Xiuying ChenEMNLP 2025
Builds on10
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- RADAR: Robust AI-Text Detection via Adversarial LearningXiaomeng Hu, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 315 citations
Related papers
- Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style AttacksJiaying Wu, Jiafeng Guo, Bryan HooiKDD 2024 · 69 citations
- FACTGUARD: Event-Centric and Commonsense-Guided Fake News DetectionJing He, Han Zhang, Yuanhui Xiao, Wei Guo et al.AAAI 2026
- Prompt-Induced Linguistic Fingerprints for LLM-Generated Fake News DetectionChi Wang, Min Gao, Zongwei Wang, Junwei Yin et al.WWW 2026 · 3 citations
- PHPFND: Detecting Fake News via Post-Hoc Processing of LLMs HallucinationJinke Ma, Jiachen Ma, Wei Zhang, Yong LiuAAAI 2026
- Adversarial Style Augmentation for Domain Generalized Urban-Scene SegmentationZhun Zhong, Yuyang Zhao, Gim Hee Lee, Nicu SebeNeurIPS 2022 · 130 citations
