Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot Detection
Chris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk, Philipp Zimmer
Abstract
Accurate bot detection is necessary for the safety and integrity of online platforms. It is also crucial for research on the influence of bots in elections, the spread of misinformation, and financial market manipulation. Platforms deploy infrastructure to flag or remove automated accounts, but their tools and data are not publicly available. Thus, the public must rely on third-party bot detection. These tools employ machine learning and often achieve near-perfect performance for classification on existing datasets, suggesting bot detection is accurate, reliable and fit for use in downstream applications. We provide evidence that this is not the case and show that high performance is attributable to limitations in dataset collection and labeling rather than sophistication of the tools. Specifically, we show that simple decision rules — shallow decision trees trained on a small number of features — achieve near-state-of-the-art performance on most available datasets and that bot detection datasets, even when combined together, do not generalize well to out-of-sample datasets. Our findings reveal that predictions are highly dependent on each dataset’s collection and labeling procedures rather than fundamental differences between bots and humans. These results have important implications for both transparency in sampling and labeling procedures and potential biases in research using existing bot detection tools for pre-processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b0dc316-04e5-4d9b-ab0c-1ad89c094154Cited by top-tier papers8
- BotMoE: Twitter Bot Detection with Community-Aware Mixtures of Modal-Specific ExpertsYuhan Liu, Zhaoxuan Tan, Heng Wang, Shangbin Feng et al.SIGIR 2023 · 54 citations
- Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AIHoujiang Liu, Anubrata Das, Alexander Boltz, Didi Zhou et al.CSCW 2024 · 23 citations
- Bots, Elections, and Controversies: Twitter Insights from Brazil's Polarised ElectionsDiogo PachecoWWW 2024 · 13 citations
- Identifying Risky Vendors in Cryptocurrency P2P MarketplacesTaro Tsuchiya, Alejandro Cuevas Villalba, Nicolas ChristinWWW 2024 · 9 citations
- How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and AnalysisHerun Wan, Minnan Luo, Zihan Ma, Guang Dai et al.EMNLP 2025 · 3 citations
Builds on3
- Scalable and Generalizable Social Bot Detection through Data SelectionKai-Cheng Yang, Onur Varol, Pik-Mai Hui, Filippo MenczerAAAI 2020 · 385 citations
- Heterogeneity-Aware Twitter Bot Detection with Relational Graph TransformersShangbin Feng, Zhaoxuan Tan, Rui Li, Minnan LuoAAAI 2022 · 138 citations
- Disagree? You Must be a Bot! How Beliefs Shape Twitter Profile PerceptionsMagdalena Wischnewski, Rebecca Bernemann, Thao Ngo, Nicole C. KrämerCHI 2021 · 22 citations
Related papers
- SoK: Machine Learning for Misinformation DetectionMadelyne Xiao, Jonathan R. MayerUSENIX Security 2025
- BotBR: Social Bot Detection with Balanced Feature Fusion and Reliability-Enhanced Graph LearningQilong Lin, Jingya ZhouSIGIR 2025 · 3 citations
- "Better Be Computer or I'm Dumb": A Large-Scale Evaluation of Humans as Audio Deepfake DetectorsKevin Warren, Tyler Tucker, Anna Crowder, Daniel Olszewski et al.CCS 2024 · 9 citations
- BIC: Twitter Bot Detection with Text-Graph Interaction and Semantic ConsistencyZhenyu Lei, Herun Wan, Wenqian Zhang, Shangbin Feng et al.ACL 2023 · 28 citations
- Bot Meets Shortcut: How Can LLMs Aid in Handling Unknown Invariance OOD Scenarios?Shiyan Zheng, Herun Wan, Minnan Luo, Junhang HuangAAAI 2026
