Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data Generation
Kung-Hsiang Huang, Kathleen R. McKeown, Preslav Nakov, Yejin Choi, Heng Ji
摘要
Despite recent advances in detecting fake news generated by neural models, their results are not readily applicable to effective detection of human-written disinformation. What limits the successful transfer between them is the sizable gap between machine-generated fake news and human-authored ones, including the notable differences in terms of style and underlying intent. With this in mind, we propose a novel framework for generating training examples that are informed by the known styles and strategies of human-authored propaganda. Specifically, we perform self-critical sequence training guided by natural language inference to ensure the validity of the generated articles, while also incorporating propaganda techniques, such as appeal to authority and loaded language. In particular, we create a new training dataset, PropaNews, with 2,256 examples, which we release for future use. Our experimental results show that fake news detectors trained on PropaNews are better at detecting human-written disinformation by 3.62–7.69% F1 score on two public datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Can LLM-Generated Misinformation Be Detected?Canyu Chen, Kai ShuICLR 2024 · 被引用 270 次
- Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style AttacksJiaying Wu, Jiafeng Guo, Bryan HooiKDD 2024 · 被引用 69 次
- FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMsXuannan Liu, Peipei Li, Huaibo Huang, Zekun Li 等ACM MM 2024 · 被引用 46 次
- Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for MisinformationMax Glockner, Yufang Hou, Iryna GurevychEMNLP 2022 · 被引用 23 次
- Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AIHoujiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou 等CSCW 2024 · 被引用 23 次
它引用的顶会 Paper8
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsQiang Ning, Hao Wu, Rujun Han, Nanyun Peng 等EMNLP 2020 · 被引用 79 次
- Fact-Enhanced Synthetic News GenerationKai Shu, Yichuan Li, Kaize Ding, Huan LiuAAAI 2021 · 被引用 39 次
相关 Paper
- Leveraging Declarative Knowledge in Text and First-Order Logic for Fine-Grained Propaganda DetectionRuize Wang, Duyu Tang, Nan Duan, Wanjun Zhong 等EMNLP 2020 · 被引用 3 次
- Detecting Cross-Modal Inconsistency to Defend Against Neural Fake NewsReuben Tan, Bryan A. Plummer, Kate SaenkoEMNLP 2020 · 被引用 9 次
- Generate First, Then Sample: Enhancing Fake News Detection with LLM-Augmented Reinforced SamplingZhao Tong, Yimeng Gu, Huidong Liu, Qiang Liu 等ACL 2025 · 被引用 14 次
- Discourse Structures Guided Fine-grained Propaganda IdentificationYuanyuan Lei, Ruihong HuangEMNLP 2023
- Embracing Domain Differences in Fake News: Cross-domain Fake News Detection using Multi-modal DataAmila Silva, Ling Luo, Shanika Karunasekera, Christopher LeckieAAAI 2021 · 被引用 170 次
