Automated Detection of Automated Traffic
Cormac Herley
摘要
We describe a method to separate abuse from legitimate traffic when we have categorical features and no labels are available. Our approach hinges on the observation that, if we could locate them, unattacked bins of a categorical feature x would allow us to estimate the benign distribution of any feature that is independent of x. We give an algorithm that finds these unattacked bins (if they exist) and show how to build an overall classifier that is suitable for very large data volumes and high levels of abuse. The approach is one-sided: our only significant assumptions about abuse are the existence of unattacked bins, and that distributions of abuse traffic do not precisely match those of benign. We evaluate on two datasets: 3 million requests from a web-server dataset and a collection of 5.1 million Twitter accounts crawled using the public API. The results confirm that the approach is successful at identifying clusters of automated behaviors. On both problems we easily outperform unsupervised methods such as Isolation Forests, and have comparable performance to Botometer on the Twitter dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- A Principled Approach for Detecting APTs in Massive Networks via Multi-Stage Causal AnalyticsJiaping Gui, Mingjie Nie, Jinyao Guo, Futai Zou 等INFOCOM 2025 · 被引用 6 次
- Online robust non-stationary estimationAbishek Sankararaman, Balakrishnan NarayanaswamyNeurIPS 2023 · 被引用 3 次
- Online Adaptive Anomaly Thresholding with Confidence SequencesSophia Huiwen Sun, Abishek Sankararaman, Balakrishnan NarayanaswamyICML 2024 · 被引用 2 次
- Estimating the Amount of Script-generated Traffic in a MixtureCormac HerleyUSENIX Security 2026
它引用的顶会 Paper3
- Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data AugmentationSteve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu 等S&P 2020 · 被引用 88 次
- Deep Entity Classification: Abusive Account Detection for Online Social NetworksTeng Xu, Gerard Goossen, Huseyin Kerem Cevahir, Sara Khodeir 等USENIX Security 2021 · 被引用 41 次
- Distinguishing Attacks from Legitimate Authentication Traffic at ScaleCormac Herley, Stuart E. SchechterNDSS 2019 · 被引用 15 次
相关 Paper
- Preventing Artificially Inflated SMS Attacks through Large-Scale Traffic InspectionJun Ho Huh, Hyejin Shin, Sunwoo Ahn, Hayoon Yi 等USENIX Security 2025
- Setting the Record Straighter on Shadow BanningErwan Le Merrer, Benoît Morgan, Gilles TrédanINFOCOM 2021 · 被引用 6 次
- FARE: Enabling Fine-grained Attack Categorization under Low-quality Labeled DataJunjie Liang, Wenbo Guo, Tongbo Luo, Vasant G. Honavar 等NDSS 2021
- Everyone's Privacy Matters! An Analysis of Privacy Leakage from Real-World Facial Images on Twitter and Associated User BehaviorsYuqi Niu, Weidong Qiu, Peng Tang, Lifan Wang 等CSCW 2025 · 被引用 3 次
- UMD: Unsupervised Model Detection for X2X Backdoor AttacksZhen Xiang, Zidi Xiong, Bo LiICML 2023 · 被引用 27 次
