Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style Attacks
Jiaying Wu, Jiafeng Guo, Bryan Hooi
摘要
It is commonly perceived that fake news and real news exhibit distinct writing styles, such as the use of sensationalist versus objective language. However, we emphasize that style-related features can also be exploited for style-based attacks. Notably, the advent of powerful Large Language Models (LLMs) has empowered malicious actors to mimic the style of trustworthy news sources, doing so swiftly, cost-effectively, and at scale. Our analysis reveals that LLM-camouflaged fake news content significantly undermines the effectiveness of state-of-the-art text-based detectors (up to 38% decrease in F1 Score), implying a severe vulnerability to stylistic variations. To address this, we introduce SheepDog, a style-robust fake news detector that prioritizes content over style in determining news veracity. SheepDog achieves this resilience through (1) LLM-empowered news reframings that inject style diversity into the training process by customizing articles to match different styles;
(2) a style-agnostic training scheme that ensures consistent veracity predictions across style-diverse reframings; and (3) content-focused veracity attributions that distill content-centric guidelines from LLMs for debunking fake news, offering supplementary cues and potential intepretability that assist veracity prediction. Extensive experiments on three real-world benchmarks demonstrate Sheep-Dog's style robustness and adaptability to various backbones. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot DetectionShangbin Feng, Herun Wan, Ningnan Wang, Zhaoxuan Tan 等ACL 2024 · 被引用 19 次
- Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation DetectionHerun Wan, Jiaying Wu, Minnan Luo, Zhi Zeng 等NeurIPS 2025 · 被引用 14 次
- Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language ModelsJiaying Wu, Fanxiao Li, Zihang Fu, Min-Yen Kan 等ICLR 2026 · 被引用 9 次
- Adversarial Style Augmentation via Large Language Model for Robust Fake News DetectionSungwon Park, Sungwon Han, Xing Xie, Jae-Gil Lee 等WWW 2025 · 被引用 9 次
- DABL: Detecting Semantic Anomalies in Business Processes Using Large Language ModelsWei Guan, Jian Cao, Jianqi Gao, Haiyan Zhao 等AAAI 2025 · 被引用 9 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
相关 Paper
- FACTGUARD: Event-Centric and Commonsense-Guided Fake News DetectionJing He, Han Zhang, Yuanhui Xiao, Wei Guo 等AAAI 2026
- Robust Fake News Detection using Large Language Models under Adversarial Sentiment AttacksSahar Tahmasebi, Eric Müller-Budack, Ralph EwerthWWW 2026 · 被引用 3 次
- PHPFND: Detecting Fake News via Post-Hoc Processing of LLMs HallucinationJinke Ma, Jiachen Ma, Wei Zhang, Yong LiuAAAI 2026
- Prompt-Induced Linguistic Fingerprints for LLM-Generated Fake News DetectionChi Wang, Min Gao, Zongwei Wang, Junwei Yin 等WWW 2026 · 被引用 3 次
- Mitigating Adversarial Attacks by Transferring LLM-generated Narrative Reasoning for Robust Fake News DetectionMengyang Chen, Lingwei Wei, Wei Zhou, Songlin HuSIGIR 2026
