The Surprising Performance of Simple Baselines for Misinformation Detection
Kellin Pelrine, Jacob Danovitch, Reihaneh Rabbany
Abstract
As social media becomes increasingly prominent in our day to day lives, it is increasingly important to detect informative content and prevent the spread of disinformation and unverified rumours. While many sophisticated and successful models have been proposed in the literature, they are often compared with older NLP baselines such as SVMs, CNNs, and LSTMs. In this paper, we examine the performance of a broad set of modern transformer-based language models and show that with basic fine-tuning, these models are competitive with and can even significantly outperform recently proposed state-of-the-art methods. We present our framework as a baseline for creating and evaluating new methods for misinformation detection. We further study a comprehensive set of benchmark datasets, and discuss potential data leakage and the need for careful design of the experiments and understanding of datasets to account for confounding variables. As an extreme case example, we show that classifying only based on the first three digits of tweet ids, which contain information on the date, gives state-of-the-art performance on a commonly used benchmark dataset for fake news detection –Twitter16. We provide a simple tool to detect this problem and suggest steps to mitigate it in future datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72dbd57b-bb1b-47aa-bd16-d8eae404bd62Cited by top-tier papers9
- Cross-modal Ambiguity Learning for Multimodal Fake News DetectionYixuan Chen, Dongsheng Li, Peng Zhang, Jie Sui et al.WWW 2022 · 325 citations
- Towards Fine-Grained Reasoning for Fake News DetectionYiqiao Jin, Xiting Wang, Ruichao Yang, Yizhou Sun et al.AAAI 2022 · 89 citations
- Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style AttacksJiaying Wu, Jiafeng Guo, Bryan HooiKDD 2024 · 69 citations
- Reinforcement Subgraph Reasoning for Fake News DetectionRuichao Yang, Xiting Wang, Yiqiao Jin, Chaozhuo Li et al.KDD 2022 · 57 citations
- DECOR: Degree-Corrected Social Graph Refinement for Fake News DetectionJiaying Wu, Bryan HooiKDD 2023 · 38 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- GCAN: Graph-aware Co-Attention Networks for Explainable Fake News Detection on Social MediaYi-Ju Lu, Cheng-Te LiACL 2020 · 387 citations
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingZihang Dai, Guokun Lai, Yiming Yang, Quoc LeNeurIPS 2020 · 273 citations
- DETERRENT: Knowledge Guided Graph Attention Network for Detecting Healthcare MisinformationLimeng Cui, Haeseung Seo, Maryam Tabar, Fenglong Ma et al.KDD 2020 · 163 citations
Related papers
- Navigating the Kaleidoscope of COVID-19 Misinformation Using Deep LearningYuanzhi Chen, Mohammad Rashedul HasanEMNLP 2021 · 4 citations
- Interpretable Rumor Detection in Microblogs by Attending to User InteractionsLing Min Serena Khoo, Hai Leong Chieu, Zhong Qian, Jing JiangAAAI 2020 · 231 citations
- Heterogeneity-Aware Twitter Bot Detection with Relational Graph TransformersShangbin Feng, Zhaoxuan Tan, Rui Li, Minnan LuoAAAI 2022 · 138 citations
- BotMoE: Twitter Bot Detection with Community-Aware Mixtures of Modal-Specific ExpertsYuhan Liu, Zhaoxuan Tan, Heng Wang, Shangbin Feng et al.SIGIR 2023 · 54 citations
- Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement MeasurementZihao Cheng, Li Zhou, Feng Jiang, Benyou Wang et al.WWW 2025 · 20 citations
