Estimating the Amount of Script-generated Traffic in a Mixture
Cormac Herley
摘要
We address the question of estimating the fraction of traffic that is bot-generated in a mixture. That is, we seek to estimate (1-α) when what we receive is α•Clean+(1-α)•Bot. This is primarily of interest when traffic is attempting to masquerade as human-generated (eg, click-fraud, inauthentic social media engagement, etc).
When at least one pair of features is independent in the clean traffic (eg, time-invariance of geographic distribution) we show that getting an upper-bound on α is equivalent to finding the rank-one matrix that maximizes a simple objective function. We give an efficient method for solving, and derive the tightness of the bound. Since the sampled version of a rank-one matrix need not be precisely rank-one, error analysis is extremely important when we have limited data. We derive error intervals for our estimates, that allow us to be confident that we find a true upper-bound.
We empirically validate our findings. First, using random rank-one, and full-rank matrices for the clean and bot distributions respectively, we verify accuracy using Monte Carlo simulations and demonstrate robustness to moderate violations of the assumptions. Second, we examine Twitter (now X) data. Twitter accounts with large follower-ship that were offered for sale on an open market-place are flagged as having > 90% bot followers, while accounts for several academic conferences and well-known researchers are flagged at < 20%. We verify accuracy on Twitter account populations of arbitrary clean/bot composition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Resident Evil: Understanding Residential IP Proxy as a Dark ServiceXianghang Mi, Xuan Feng, Xiaojing Liao, Baojun Liu 等S&P 2019 · 被引用 80 次
- Deep Entity Classification: Abusive Account Detection for Online Social NetworksTeng Xu, Gerard Goossen, Huseyin Kerem Cevahir, Sara Khodeir 等USENIX Security 2021 · 被引用 41 次
- Automated Detection of Automated TrafficCormac HerleyUSENIX Security 2022
相关 Paper
- BotMoE: Twitter Bot Detection with Community-Aware Mixtures of Modal-Specific ExpertsYuhan Liu, Zhaoxuan Tan, Heng Wang, Shangbin Feng 等SIGIR 2023 · 被引用 54 次
- Scalable and Generalizable Social Bot Detection through Data SelectionKai-Cheng Yang, Onur Varol, Pik-Mai Hui, Filippo MenczerAAAI 2020 · 被引用 385 次
- Beyond Bot Detection: Combating Fraudulent Online Survey Takers✱Ziyi Zhang, Shuofei Zhu, Jaron Mink, Aiping Xiong 等WWW 2022 · 被引用 51 次
- Disagree? You Must be a Bot! How Beliefs Shape Twitter Profile PerceptionsMagdalena Wischnewski, Rebecca Bernemann, Thao Ngo, Nicole C. KrämerCHI 2021 · 被引用 22 次
- Robust Graph Matching when Nodes are CorruptTaha Ameen, Bruce E. HajekICML 2024 · 被引用 7 次
