Low-Quality Training Data Only? A Robust Framework for Detecting Encrypted Malicious Network Traffic
Yuqi Qing, Qilei Yin, Xinhao Deng, Yihao Chen, Zhuotao Liu, Kun Sun, Ke Xu, Jia Zhang, Qi Li
摘要
Machine learning (ML) is promising in accurately detecting malicious flows in encrypted network traffic; however, it is challenging to collect a training dataset that contains a sufficient amount of encrypted malicious data with correct labels. When ML models are trained with low-quality training data, they suffer degraded performance. In this paper, we aim at addressing a real-world low-quality training dataset problem, namely, detecting encrypted malicious traffic generated by continuously evolving malware. We develop RAPIER that fully utilizes different distributions of normal and malicious traffic data in the feature space, where normal data is tightly distributed in a certain area and the malicious data is scattered over the entire feature space to augment training data for model training. RAPIER includes two pre-processing modules to convert traffic into feature vectors and correct label noises. We evaluate our system on two public datasets and one combined dataset. With 1000 samples and 45% noises from each dataset, our system achieves the F1 scores of 0.770, 0.776, and 0.855, respectively, achieving average improvements of 352.6%, 284.3%, and 214.9% over the existing methods, respectively. Furthermore, We evaluate RAPIER with a real-world dataset obtained from a security enterprise. RAPIER effectively achieves encrypted malicious traffic detection with the best F1 score of 0.773 and improves the F1 score of existing methods by an average of 272.5%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Brain-on-Switch: Towards Advanced Intelligent Network Data Plane via NN-Driven Traffic Analysis at Line-SpeedJinzhu Yan, Haotian Xu, Zhuotao Liu, Qi Li 等NSDI 2024 · 被引用 60 次
- Point Cloud Analysis for ML-Based Malicious Traffic Detection: Reducing Majorities of False Positive AlarmsChuanpu Fu, Qi Li, Ke Xu, Jianping WuCCS 2023 · 被引用 30 次
- Robust and Reliable Early-Stage Website Fingerprinting Attacks via Spatial-Temporal Distribution AnalysisXinhao Deng, Qi Li, Ke XuCCS 2024 · 被引用 22 次
- Detecting Tunneled Flooding Traffic via Deep Semantic Analysis of Packet Length PatternsChuanpu Fu, Qi Li, Meng Shen, Ke XuCCS 2024 · 被引用 13 次
- FENIX: Enabling In-Network DNN Inference with FPGA-Enhanced Programmable SwitchesXiangyu Gao, Tong Li, Yinchao Zhang, Ziqiang Wang 等NSDI 2026 · 被引用 12 次
它引用的顶会 Paper28
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Kitsune: An Ensemble of Autoencoders for Online Network Intrusion DetectionYisroel Mirsky, Tomer Doitshman, Yuval Elovici, Asaf ShabtaiNDSS 2018 · 被引用 945 次
- Deep Fingerprinting: Undermining Website Fingerprinting Defenses with Deep LearningPayap Sirinam, Mohsen Imani, Marc Juarez, Matthew WrightCCS 2018 · 被引用 632 次
- ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic ClassificationXinjie Lin, Gang Xiong, Gaopeng Gou, Zhen Li 等WWW 2022 · 被引用 490 次
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder 等USENIX Security 2019 · 被引用 441 次
相关 Paper
- Detecting Unknown Encrypted Malicious Traffic in Real Time via Flow Interaction Graph AnalysisChuanpu Fu, Qi Li, Ke XuNDSS 2023
- Wedjat: Detecting Sophisticated Evasion Attacks via Real-time Causal AnalysisLi Gao, Chuanpu Fu, Xinhao Deng, Ke Xu 等KDD 2025 · 被引用 2 次
- Realtime Robust Malicious Traffic Detection via Frequency Domain AnalysisChuanpu Fu, Qi Li, Meng Shen, Ke XuCCS 2021 · 被引用 194 次
- CND-IDS: Continual Novelty Detection for Intrusion Detection SystemsSean Fuhrman, Onat Güngör, Tajana RosingDAC 2025 · 被引用 9 次
- Training Robust Classifiers for Classifying Encrypted Traffic under Dynamic Network ConditionsYuqi Qing, Qilei Yin, Xinhao Deng, Xiaoli Zhang 等CCS 2025
