What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot Detection
Shangbin Feng, Herun Wan, Ningnan Wang, Zhaoxuan Tan, Minnan Luo, Yulia Tsvetkov
Abstract
Social media bot detection has always been an arms race between advancements in machine learning bot detectors and adversarial bot strategies to evade detection. In this work, we bring the arms race to the next level by investigating the opportunities and risks of state-of-the-art large language models (LLMs) in social bot detection. To investigate the opportunities, we design novel LLM-based bot detectors by proposing a mixture-of-heterogeneous-experts framework to divide and conquer diverse user information modalities. To illuminate the risks, we explore the possibility of LLM-guided manipulation of user textual and structured information to evade detection. Extensive experiments with three LLMs on two datasets demonstrate that instruction tuning on merely 1,000 annotated examples produces specialized LLMs that outperform state-of-the-art bot detection baselines by up to 9.1% on both datasets. On the other hand, LLM-guided manipulation strategies could significantly bring down the performance of existing bot detectors by up to 29.6% and harm the calibration and reliability of bot detection systems. Ultimately, this works identifies LLMs as the new frontier of social bot detection research. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- BotSim: LLM-Powered Malicious Social Botnet SimulationBoyu Qiao, Kun Li, Wei Zhou, Shilong Li et al.AAAI 2025 · 23 citations
- How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and AnalysisHerun Wan, Minnan Luo, Zihan Ma, Guang Dai et al.EMNLP 2025 · 3 citations
- TBTrackerX: Fantastic Trigger Bots and Where to Find Malicious Campaigns on XMohammad Majid Akhtar, Rahat Masood, Muhammad Ikram, Salil S. KanhereNDSS 2026 · 1 citation
- GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language ModelYunhe Pang, Bo Chen, Fanjin Zhang, Yanghui Rao et al.KDD 2025
- Exploring and Distilling Multi-Dimensional Clues for Interpretable Social Bot DetectionYi Han, Haiqi Lu, Lizi Liao, Shuhan Zhou et al.ACL 2026
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
Related papers
- Enhancing LLM-Based Social Bot via an Adversarial Learning FrameworkFanqi Kong, Xiaoyuan Zhang, Xinyu Chen, Yaodong Yang et al.EMNLP 2025 · 1 citation
- Large Language Model (LLM)-driven Adversarial Social Influences in Online Information Spread: Risks and InterventionsZhuoran Lu, Gionnieve Lim, Ming YinCHI 2026
- On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMsHerun Wan, Minnan Luo, Zhixiong Su, Guang Dai et al.ACL 2025 · 5 citations
- Bot Meets Shortcut: How Can LLMs Aid in Handling Unknown Invariance OOD Scenarios?Shiyan Zheng, Herun Wan, Minnan Luo, Junhang HuangAAAI 2026
- Language Model Detectors Are Easily Optimized AgainstCharlotte Nicks, Eric Mitchell, Rafael Rafailov, Archit Sharma et al.ICLR 2024 · 18 citations
