USENIX Security2022Top-tier venue
Automated Detection of Automated Traffic
Cormac Herley
Abstract
We describe a method to separate abuse from legitimate traffic when we have categorical features and no labels are available. Our approach hinges on the observation that, if we could locate them, unattacked bins of a categorical feature x would allow us to estimate the benign distribution of any feature that is independent of x. We give an algorithm that finds these unattacked bins (if they exist) and show how to build an overall classifier that is suitable for very large data volumes and high levels of abuse. The approach is one-sided: our only significant assumptions about abuse are the existence of unattacked bins, and that distributions of abuse traffic do not precisely match those of benign. We evaluate on two datasets: 3 million requests from a web-server dataset and a collection of 5.1 million Twitter accounts crawled using the public API. The results confirm that the approach is successful at identifying clusters of automated behaviors. On both problems we easily outperform unsupervised methods such as Isolation Forests, and have comparable performance to Botometer on the Twitter dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- A Principled Approach for Detecting APTs in Massive Networks via Multi-Stage Causal AnalyticsJiaping Gui, Mingjie Nie, Jinyao Guo, Futai Zou et al.INFOCOM 2025 · 6 citations
- Online robust non-stationary estimationAbishek Sankararaman, Balakrishnan NarayanaswamyNeurIPS 2023 · 3 citations
- Online Adaptive Anomaly Thresholding with Confidence SequencesSophia Huiwen Sun, Abishek Sankararaman, Balakrishnan NarayanaswamyICML 2024 · 2 citations
- Estimating the Amount of Script-generated Traffic in a MixtureCormac HerleyUSENIX Security 2026
Builds on3
- Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data AugmentationSteve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu et al.S&P 2020 · 88 citations
- Deep Entity Classification: Abusive Account Detection for Online Social NetworksTeng Xu, Gerard Goossen, Huseyin Kerem Cevahir, Sara Khodeir et al.USENIX Security 2021 · 41 citations
- Distinguishing Attacks from Legitimate Authentication Traffic at ScaleCormac Herley, Stuart E. SchechterNDSS 2019 · 15 citations
Related papers
- Preventing Artificially Inflated SMS Attacks through Large-Scale Traffic InspectionJun Ho Huh, Hyejin Shin, Sunwoo Ahn, Hayoon Yi et al.USENIX Security 2025
- Setting the Record Straighter on Shadow BanningErwan Le Merrer, Benoît Morgan, Gilles TrédanINFOCOM 2021 · 6 citations
- FARE: Enabling Fine-grained Attack Categorization under Low-quality Labeled DataJunjie Liang, Wenbo Guo, Tongbo Luo, Vasant G. Honavar et al.NDSS 2021
- Everyone's Privacy Matters! An Analysis of Privacy Leakage from Real-World Facial Images on Twitter and Associated User BehaviorsYuqi Niu, Weidong Qiu, Peng Tang, Lifan Wang et al.CSCW 2025 · 3 citations
- UMD: Unsupervised Model Detection for X2X Backdoor AttacksZhen Xiang, Zidi Xiong, Bo LiICML 2023 · 27 citations
