USENIX Security2025Top-tier venue
SoK: Machine Learning for Misinformation Detection
Madelyne Xiao, Jonathan R. Mayer
Abstract
We examine the disconnect between scholarship and practice in applying machine learning to trust and safety problems, using misinformation detection as a case study. We survey literature on automated detection of misinformation across a corpus of 248 well-cited papers in the field. We then examine subsets of papers for data and code availability, design missteps, reproducibility, and generalizability. Our paper corpus includes published work in security, natural language processing, and computational social science. Across these disparate disciplines, we identify common errors in dataset and method design. In general, detection tasks are often meaningfully distinct from the challenges that online services actually face. Datasets and model evaluation are often non-representative of real-world contexts, and evaluation frequently is not independent of model training. We demonstrate the limitations of current detection methods in a series of three representative replication studies. Based on the results of these analyses and our literature survey, we conclude that the current state-of-the-art in fully-automated misinformation detection has limited efficacy in detecting human-generated misinformation. We offer recommendations for evaluating applications of machine learning to trust and safety problems and recommend future directions for research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0c9fe8a-2434-4ce5-87cd-96fb5979c877Cited by top-tier papers2
- Chameleon Channels: Measuring YouTube Accounts Repurposed for Deception and ProfitAlejandro Cuevas, Manoel Horta Ribeiro, Nicolas ChristinUSENIX Security 2026
- Cross-National Information Attacks: A Two-Decade Analysis of Troll Behavior in KoreaJaehong Kim, Hyeonseung Kim, Jiseon Kim, Alice Oh et al.USENIX Security 2026
Builds on12
- HOLMES: Real-Time APT Detection through Correlation of Suspicious Information FlowsSadegh Momeni Milajerdi, Rigel Gjomemo, Birhanu Eshete, R. Sekar et al.S&P 2019 · 550 citations
- Embracing Domain Differences in Fake News: Cross-domain Fake News Detection using Multi-modal DataAmila Silva, Ling Luo, Shanika Karunasekera, Christopher LeckieAAAI 2021 · 170 citations
- SoK: A Framework for Unifying At-Risk User ResearchNoel Warford, Tara Matthews, Kaitlyn Yang, Omer Akgul et al.S&P 2022 · 101 citations
- Detecting Fake Accounts in Online Social Networks at the Time of RegistrationsDong Yuan, Yuanli Miao, Neil Zhenqiang Gong, Zheng Yang et al.CCS 2019 · 86 citations
- TrollMagnifier: Detecting State-Sponsored Troll Accounts on RedditMohammad Hammas Saeed, Shiza Ali, Jeremy Blackburn, Emiliano De Cristofaro et al.S&P 2022 · 40 citations
Related papers
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Can LLM-Generated Misinformation Be Detected?Canyu Chen, Kai ShuICLR 2024 · 270 citations
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller et al.NDSS 2026 · 17 citations
- Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot DetectionChris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk et al.WWW 2023 · 38 citations
- How does Misinformation Affect Large Language Model Behaviors and Preferences?Miao Peng, Nuo Chen, Jianheng Tang, Jia LiACL 2025 · 2 citations
