USENIX Security2026Top-tier venue
Estimating the Amount of Script-generated Traffic in a Mixture
Cormac Herley
Abstract
We address the question of estimating the fraction of traffic that is bot-generated in a mixture. That is, we seek to estimate (1-α) when what we receive is α•Clean+(1-α)•Bot. This is primarily of interest when traffic is attempting to masquerade as human-generated (eg, click-fraud, inauthentic social media engagement, etc).
When at least one pair of features is independent in the clean traffic (eg, time-invariance of geographic distribution) we show that getting an upper-bound on α is equivalent to finding the rank-one matrix that maximizes a simple objective function. We give an efficient method for solving, and derive the tightness of the bound. Since the sampled version of a rank-one matrix need not be precisely rank-one, error analysis is extremely important when we have limited data. We derive error intervals for our estimates, that allow us to be confident that we find a true upper-bound.
We empirically validate our findings. First, using random rank-one, and full-rank matrices for the clean and bot distributions respectively, we verify accuracy using Monte Carlo simulations and demonstrate robustness to moderate violations of the assumptions. Second, we examine Twitter (now X) data. Twitter accounts with large follower-ship that were offered for sale on an open market-place are flagged as having > 90% bot followers, while accounts for several academic conferences and well-known researchers are flagged at < 20%. We verify accuracy on Twitter account populations of arbitrary clean/bot composition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Resident Evil: Understanding Residential IP Proxy as a Dark ServiceXianghang Mi, Xuan Feng, Xiaojing Liao, Baojun Liu et al.S&P 2019 · 80 citations
- Deep Entity Classification: Abusive Account Detection for Online Social NetworksTeng Xu, Gerard Goossen, Huseyin Kerem Cevahir, Sara Khodeir et al.USENIX Security 2021 · 41 citations
- Automated Detection of Automated TrafficCormac HerleyUSENIX Security 2022
Related papers
- BotMoE: Twitter Bot Detection with Community-Aware Mixtures of Modal-Specific ExpertsYuhan Liu, Zhaoxuan Tan, Heng Wang, Shangbin Feng et al.SIGIR 2023 · 54 citations
- Scalable and Generalizable Social Bot Detection through Data SelectionKai-Cheng Yang, Onur Varol, Pik-Mai Hui, Filippo MenczerAAAI 2020 · 385 citations
- Beyond Bot Detection: Combating Fraudulent Online Survey Takers✱Ziyi Zhang, Shuofei Zhu, Jaron Mink, Aiping Xiong et al.WWW 2022 · 51 citations
- Disagree? You Must be a Bot! How Beliefs Shape Twitter Profile PerceptionsMagdalena Wischnewski, Rebecca Bernemann, Thao Ngo, Nicole C. KrämerCHI 2021 · 22 citations
- Robust Graph Matching when Nodes are CorruptTaha Ameen, Bruce E. HajekICML 2024 · 7 citations
