The Surprising Performance of Simple Baselines for Misinformation Detection
Kellin Pelrine, Jacob Danovitch, Reihaneh Rabbany
摘要
As social media becomes increasingly prominent in our day to day lives, it is increasingly important to detect informative content and prevent the spread of disinformation and unverified rumours. While many sophisticated and successful models have been proposed in the literature, they are often compared with older NLP baselines such as SVMs, CNNs, and LSTMs. In this paper, we examine the performance of a broad set of modern transformer-based language models and show that with basic fine-tuning, these models are competitive with and can even significantly outperform recently proposed state-of-the-art methods. We present our framework as a baseline for creating and evaluating new methods for misinformation detection. We further study a comprehensive set of benchmark datasets, and discuss potential data leakage and the need for careful design of the experiments and understanding of datasets to account for confounding variables. As an extreme case example, we show that classifying only based on the first three digits of tweet ids, which contain information on the date, gives state-of-the-art performance on a commonly used benchmark dataset for fake news detection –Twitter16. We provide a simple tool to detect this problem and suggest steps to mitigate it in future datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Cross-modal Ambiguity Learning for Multimodal Fake News DetectionYixuan Chen, Dongsheng Li, Peng Zhang, Jie Sui 等WWW 2022 · 被引用 325 次
- Towards Fine-Grained Reasoning for Fake News DetectionYiqiao Jin, Xiting Wang, Ruichao Yang, Yizhou Sun 等AAAI 2022 · 被引用 89 次
- Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style AttacksJiaying Wu, Jiafeng Guo, Bryan HooiKDD 2024 · 被引用 69 次
- Reinforcement Subgraph Reasoning for Fake News DetectionRuichao Yang, Xiting Wang, Yiqiao Jin, Chaozhuo Li 等KDD 2022 · 被引用 57 次
- DECOR: Degree-Corrected Social Graph Refinement for Fake News DetectionJiaying Wu, Bryan HooiKDD 2023 · 被引用 38 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- GCAN: Graph-aware Co-Attention Networks for Explainable Fake News Detection on Social MediaYi-Ju Lu, Cheng-Te LiACL 2020 · 被引用 387 次
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingZihang Dai, Guokun Lai, Yiming Yang, Quoc LeNeurIPS 2020 · 被引用 273 次
- DETERRENT: Knowledge Guided Graph Attention Network for Detecting Healthcare MisinformationLimeng Cui, Haeseung Seo, Maryam Tabar, Fenglong Ma 等KDD 2020 · 被引用 163 次
相关 Paper
- Navigating the Kaleidoscope of COVID-19 Misinformation Using Deep LearningYuanzhi Chen, Mohammad Rashedul HasanEMNLP 2021 · 被引用 4 次
- Interpretable Rumor Detection in Microblogs by Attending to User InteractionsLing Min Serena Khoo, Hai Leong Chieu, Zhong Qian, Jing JiangAAAI 2020 · 被引用 231 次
- Heterogeneity-Aware Twitter Bot Detection with Relational Graph TransformersShangbin Feng, Zhaoxuan Tan, Rui Li, Minnan LuoAAAI 2022 · 被引用 138 次
- BotMoE: Twitter Bot Detection with Community-Aware Mixtures of Modal-Specific ExpertsYuhan Liu, Zhaoxuan Tan, Heng Wang, Shangbin Feng 等SIGIR 2023 · 被引用 54 次
- Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement MeasurementZihao Cheng, Li Zhou, Feng Jiang, Benyou Wang 等WWW 2025 · 被引用 20 次
