Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style Attacks
Jiaying Wu, Jiafeng Guo, Bryan Hooi
Abstract
It is commonly perceived that fake news and real news exhibit distinct writing styles, such as the use of sensationalist versus objective language. However, we emphasize that style-related features can also be exploited for style-based attacks. Notably, the advent of powerful Large Language Models (LLMs) has empowered malicious actors to mimic the style of trustworthy news sources, doing so swiftly, cost-effectively, and at scale. Our analysis reveals that LLM-camouflaged fake news content significantly undermines the effectiveness of state-of-the-art text-based detectors (up to 38% decrease in F1 Score), implying a severe vulnerability to stylistic variations. To address this, we introduce SheepDog, a style-robust fake news detector that prioritizes content over style in determining news veracity. SheepDog achieves this resilience through (1) LLM-empowered news reframings that inject style diversity into the training process by customizing articles to match different styles;
(2) a style-agnostic training scheme that ensures consistent veracity predictions across style-diverse reframings; and (3) content-focused veracity attributions that distill content-centric guidelines from LLMs for debunking fake news, offering supplementary cues and potential intepretability that assist veracity prediction. Extensive experiments on three real-world benchmarks demonstrate Sheep-Dog's style robustness and adaptability to various backbones. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79b6e08f-a5f6-40cf-ac7f-13f7d5e8474bCited by top-tier papers28
- What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot DetectionShangbin Feng, Herun Wan, Ningnan Wang, Zhaoxuan Tan et al.ACL 2024 · 19 citations
- Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation DetectionHerun Wan, Jiaying Wu, Minnan Luo, Zhi Zeng et al.NeurIPS 2025 · 14 citations
- Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language ModelsJiaying Wu, Fanxiao Li, Zihang Fu, Min-Yen Kan et al.ICLR 2026 · 9 citations
- Adversarial Style Augmentation via Large Language Model for Robust Fake News DetectionSungwon Park, Sungwon Han, Xing Xie, Jae-Gil Lee et al.WWW 2025 · 9 citations
- DABL: Detecting Semantic Anomalies in Business Processes Using Large Language ModelsWei Guan, Jian Cao, Jianqi Gao, Haiyan Zhao et al.AAAI 2025 · 9 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
Related papers
- FACTGUARD: Event-Centric and Commonsense-Guided Fake News DetectionJing He, Han Zhang, Yuanhui Xiao, Wei Guo et al.AAAI 2026
- Robust Fake News Detection using Large Language Models under Adversarial Sentiment AttacksSahar Tahmasebi, Eric Müller-Budack, Ralph EwerthWWW 2026 · 3 citations
- PHPFND: Detecting Fake News via Post-Hoc Processing of LLMs HallucinationJinke Ma, Jiachen Ma, Wei Zhang, Yong LiuAAAI 2026
- Prompt-Induced Linguistic Fingerprints for LLM-Generated Fake News DetectionChi Wang, Min Gao, Zongwei Wang, Junwei Yin et al.WWW 2026 · 3 citations
- Mitigating Adversarial Attacks by Transferring LLM-generated Narrative Reasoning for Robust Fake News DetectionMengyang Chen, Lingwei Wei, Wei Zhou, Songlin HuSIGIR 2026
