Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data Augmentation
Steve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu, Sonal Oswal, Gang Wang, Bimal Viswanath
Abstract
Machine learning has been widely applied to building security applications. However, many machine learning models require the continuous supply of representative labeled data for training, which limits the models’ usefulness in practice. In this paper, we use bot detection as an example to explore the use of data synthesis to address this problem. We collected the network traffic from 3 online services in three different months within a year (23 million network requests). We develop a stream-based feature encoding scheme to support machine learning models for detecting advanced bots. The key novelty is that our model detects bots with extremely limited labeled data. We propose a data synthesis method to synthesize unseen (or future) bot behavior distributions. The synthesis method is distribution-aware, using two different generators in a Generative Adversarial Network to synthesize data for the clustered regions and the outlier regions in the feature space. We evaluate this idea and show our method can train a model that outperforms existing methods with only 1% of the labeled data. We show that data synthesis also improves the model’s sustainability over time and speeds up the retraining. Finally, we compare data synthesis and adversarial retraining and show they can work complementary with each other to improve the model generalizability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers30
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi et al.USENIX Security 2021 · 241 citations
- Realtime Robust Malicious Traffic Detection via Frequency Domain AnalysisChuanpu Fu, Qi Li, Meng Shen, Ke XuCCS 2021 · 194 citations
- Brain-on-Switch: Towards Advanced Intelligent Network Data Plane via NN-Driven Traffic Analysis at Line-SpeedJinzhu Yan, Haotian Xu, Zhuotao Liu, Qi Li et al.NSDI 2024 · 60 citations
- Good Bot, Bad Bot: Characterizing Automated Browsing ActivityXigao Li, Babak Amin Azad, Amir Rahmati, Nick NikiforakisS&P 2021 · 45 citations
- Point Cloud Analysis for ML-Based Malicious Traffic Detection: Reducing Majorities of False Positive AlarmsChuanpu Fu, Qi Li, Ke Xu, Jianping WuCCS 2023 · 30 citations
Builds on13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Kitsune: An Ensemble of Autoencoders for Online Network Intrusion DetectionYisroel Mirsky, Tomer Doitshman, Yuval Elovici, Asaf ShabtaiNDSS 2018 · 945 citations
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder et al.USENIX Security 2019 · 441 citations
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang et al.USENIX Security 2017 · 325 citations
- When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning AttacksOctavian Suciu, Radu Marginean, Yigitcan Kaya, Hal Daumé III et al.USENIX Security 2018 · 321 citations
Related papers
- CND-IDS: Continual Novelty Detection for Intrusion Detection SystemsSean Fuhrman, Onat Güngör, Tajana RosingDAC 2025 · 9 citations
- Frequency-Domain Mixing Data Augmentation for Malicious Traffic DetectionYuhao Yan, Bo Lang, Xiangyu LiCCS 2026
- Knowledge Enhanced GAN for IoT Traffic GenerationShuodi Hui, Huandong Wang, Zhenhua Wang, Xinghao Yang et al.WWW 2022 · 54 citations
- Learning from Few Samples: A Novel Approach for High-Quality Malcode GenerationHaijian Ma, Daizong Liu, Xiaowen Cai, Pan Zhou et al.EMNLP 2025
- AdvTG: An Adversarial Traffic Generation Framework to Deceive DL-Based Malicious Traffic Detection ModelsPeishuai Sun, Xiaochun Yun, Shuhao Li, Tao Yin et al.WWW 2025 · 5 citations
