Scalable and Generalizable Social Bot Detection through Data Selection
Kai-Cheng Yang, Onur Varol, Pik-Mai Hui, Filippo Menczer
Abstract
Efficient and reliable social bot classification is crucial for detecting information manipulation on social media. Despite rapid development, state-of-the-art bot detection models still face generalization and scalability challenges, which greatly limit their applications. In this paper we propose a framework that uses minimal account metadata, enabling efficient analysis that scales up to handle the full stream of public tweets of Twitter in real time. To ensure model accuracy, we build a rich collection of labeled datasets for training and validation. We deploy a strict validation system so that model performance on unseen datasets is also optimized, in addition to traditional cross-validation. We find that strategically selecting a subset of training data yields better model accuracy and generalization than exhaustively training on all available data. Thanks to the simplicity of the proposed model, its logic can be interpreted to provide insights into social bot characteristics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24677913-031d-4a49-a150-89b7e3c012aeCited by top-tier papers24
- Heterogeneity-Aware Twitter Bot Detection with Relational Graph TransformersShangbin Feng, Zhaoxuan Tan, Rui Li, Minnan LuoAAAI 2022 · 138 citations
- BotMoE: Twitter Bot Detection with Community-Aware Mixtures of Modal-Specific ExpertsYuhan Liu, Zhaoxuan Tan, Heng Wang, Shangbin Feng et al.SIGIR 2023 · 54 citations
- FedACK: Federated Adversarial Contrastive Knowledge Distillation for Cross-Lingual and Cross-Model Social Bot DetectionYingguang Yang, Renyu Yang, Hao Peng, Yangyang Li et al.WWW 2023 · 39 citations
- Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot DetectionChris Hays, Zachary Schutzman, Manish Raghavan, Erin Walk et al.WWW 2023 · 38 citations
- BIC: Twitter Bot Detection with Text-Graph Interaction and Semantic ConsistencyZhenyu Lei, Herun Wan, Wenqian Zhang, Shangbin Feng et al.ACL 2023 · 28 citations
Related papers
- ETS-MM: A Multi-Modal Social Bot Detection Model Based on Enhanced Textual Semantic RepresentationWei Li, Jiawen Deng, Jiali You, Yuanyuan He et al.WWW 2025 · 10 citations
- BSG4Bot:Efficient Bot Detection Based on Biased Heterogeneous SubgraphsHao Miao, Zida Liu, Jun GaoICDE 2025 · 1 citation
- OTPCL: Optimal Transport Driven Pseudo-Labeling with Contrastive Learning for Social Bot DetectionRuixuan Xu, Mengting Hu, Xinqi Yang, Ming Jiang et al.KDD 2026
- How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and AnalysisHerun Wan, Minnan Luo, Zihan Ma, Guang Dai et al.EMNLP 2025 · 3 citations
- Adversarial Socialbots Modeling Based on Structural Information PrinciplesXianghua Zeng, Hao Peng, Angsheng LiAAAI 2024 · 27 citations
