Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data Generation
Kung-Hsiang Huang, Kathleen R. McKeown, Preslav Nakov, Yejin Choi, Heng Ji
Abstract
Despite recent advances in detecting fake news generated by neural models, their results are not readily applicable to effective detection of human-written disinformation. What limits the successful transfer between them is the sizable gap between machine-generated fake news and human-authored ones, including the notable differences in terms of style and underlying intent. With this in mind, we propose a novel framework for generating training examples that are informed by the known styles and strategies of human-authored propaganda. Specifically, we perform self-critical sequence training guided by natural language inference to ensure the validity of the generated articles, while also incorporating propaganda techniques, such as appeal to authority and loaded language. In particular, we create a new training dataset, PropaNews, with 2,256 examples, which we release for future use. Our experimental results show that fake news detectors trained on PropaNews are better at detecting human-written disinformation by 3.62–7.69% F1 score on two public datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93176b6a-61de-4a83-9daa-b3ff8b955a18Cited by top-tier papers25
- Can LLM-Generated Misinformation Be Detected?Canyu Chen, Kai ShuICLR 2024 · 270 citations
- Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style AttacksJiaying Wu, Jiafeng Guo, Bryan HooiKDD 2024 · 69 citations
- FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMsXuannan Liu, Peipei Li, Huaibo Huang, Zekun Li et al.ACM MM 2024 · 46 citations
- Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for MisinformationMax Glockner, Yufang Hou, Iryna GurevychEMNLP 2022 · 23 citations
- Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AIHoujiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou et al.CSCW 2024 · 23 citations
Builds on8
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsQiang Ning, Hao Wu, Rujun Han, Nanyun Peng et al.EMNLP 2020 · 79 citations
- Fact-Enhanced Synthetic News GenerationKai Shu, Yichuan Li, Kaize Ding, Huan LiuAAAI 2021 · 39 citations
Related papers
- Leveraging Declarative Knowledge in Text and First-Order Logic for Fine-Grained Propaganda DetectionRuize Wang, Duyu Tang, Nan Duan, Wanjun Zhong et al.EMNLP 2020 · 3 citations
- Detecting Cross-Modal Inconsistency to Defend Against Neural Fake NewsReuben Tan, Bryan A. Plummer, Kate SaenkoEMNLP 2020 · 9 citations
- Generate First, Then Sample: Enhancing Fake News Detection with LLM-Augmented Reinforced SamplingZhao Tong, Yimeng Gu, Huidong Liu, Qiang Liu et al.ACL 2025 · 14 citations
- Discourse Structures Guided Fine-grained Propaganda IdentificationYuanyuan Lei, Ruihong HuangEMNLP 2023
- Embracing Domain Differences in Fake News: Cross-domain Fake News Detection using Multi-modal DataAmila Silva, Ling Luo, Shanika Karunasekera, Christopher LeckieAAAI 2021 · 170 citations
