Frequency-Domain Mixing Data Augmentation for Malicious Traffic Detection
Yuhao Yan, Bo Lang, Xiangyu Li
Abstract
The strong dynamics of network traffic often force malicious traffic detection models to handle out-of-distribution data. Typically, deep learning-based malicious traffic detection models require a large amount of high-quality training data. However, owing to challenges such as high labeling difficulty and resource consumption, existing datasets often suffer from insufficient diversity and fail to capture evolving traffic patterns, leading to poor out-of-distribution generalization ability of the trained models. Data augmentation has been widely adopted to improve data diversity and model generalization. Recently, frequency-domain mixing augmentation has shown promising performance because it effectively perturbs data while preserving key structural information. This approach shows potential for enhancing malicious traffic detection models. However, existing studies lack theoretical interpretation of the mixing mechanism, and do not adapt to the characteristics of network traffic. In this paper, we first conduct a theoretical analysis of the current frequency-domain mixing method, revealing its underlying principles and limitations. We further propose an improved frequency-domain mixing-based data augmentation method for network traffic data, which enhances the diversity of sequence features in network traffic and improves the out-of-distribution generalization of malicious traffic detection models. Extensive experiments on multiple artificial and real-world datasets demonstrate that our method substantially improves detection performance across diverse network environments and outperforms other data augmentation approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Kitsune: An Ensemble of Autoencoders for Online Network Intrusion DetectionYisroel Mirsky, Tomer Doitshman, Yuval Elovici, Asaf ShabtaiNDSS 2018 · 945 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Deep Fingerprinting: Undermining Website Fingerprinting Defenses with Deep LearningPayap Sirinam, Mohsen Imani, Marc Juarez, Matthew WrightCCS 2018 · 632 citations
- Website Fingerprinting at Internet ScaleAndriy Panchenko, Fabian Lanze, Jan Pennekamp, Thomas Engel et al.NDSS 2016 · 625 citations
- ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic ClassificationXinjie Lin, Gang Xiong, Gaopeng Gou, Zhen Li et al.WWW 2022 · 490 citations
Related papers
- Domain Generalization with Vital Phase AugmentationIngyun Lee, Wooju Lee, Hyun MyungAAAI 2024 · 12 citations
- Out-Of-Distribution Detection with Diversification (Provably)Haiyun Yao, Zongbo Han, Huazhu Fu, Xi Peng et al.NeurIPS 2024 · 9 citations
- Realtime Robust Malicious Traffic Detection via Frequency Domain AnalysisChuanpu Fu, Qi Li, Meng Shen, Ke XuCCS 2021 · 194 citations
- Robust Image Denoising Through Adversarial Frequency MixupDonghun Ryou, Inju Ha, Hyewon Yoo, Dongwan Kim et al.CVPR 2024 · 13 citations
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency DebiasingHossein Kashiani, Niloufar Alipour Talemi, Fatemeh AfghahCVPR 2025
