Training with Only 1.0 ‰ Samples: Malicious Traffic Detection via Cross-Modality Feature Fusion
Chuanpu Fu, Qi Li, Elisa Bertino, Ke Xu
Abstract
Machine Learning (ML) based malicious traffic detection systems can accurately recognize unseen network attacks by learning from large-scale traffic datasets. However, deploying such systems across multiple networks involves substantial efforts to construct large training datasets for each network. This paper addresses the issue of training with minimal datasets, that is, achieving accurate malicious traffic detection by learning a small portion of traffic in entirely new network environments, thereby eliminating prohibitive labor costs associated with traffic dataset construction. We develop tFusion to effectively extract information from limited datasets by treating network traffic data as multimodal data, comprising features from multiple sensory modalities of packets, flows, and hosts. In particular, we design a dedicated crossmodal attention model that fuses fine-grained per-packet sequential features with coarse-grained per-flow and per-host statistical features, to synthesize correlations among the different granularities of traffic features. Moreover, we design a topology-driven contrastive learning approach that pre- trains the models while reducing topology-related biases, which allows tFusion to achieve generic detection across various networks. We deploy tFusion in an institutional network and measure its performance over five days. tFusion requires human experts to label only 1.0 ‰ traffic, yet it achieves 99.82% accuracy when detecting various attacks. Meanwhile, it outperforms 14 existing methods by improving over 12.76% accuracy on 11 existing datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- TriFusion-IDS: A Multimodal Graph-Tabular-Text Contrastive Framework for Cross-Dataset Intrusion DetectionQinxin Zhao, Sheng ZhongAAAI 2026
- ContraMTD: An Unsupervised Malicious Network Traffic Detection Method based on Contrastive LearningXueying Han, Susu Cui, Jian Qin, Song Liu et al.WWW 2024 · 25 citations
- MM4flow: A Pre-trained Multi-modal Model for Versatile Network Traffic AnalysisLuming Yang, Lin Liu, Junjie Huang, Zhuotao Liu et al.CCS 2025
- BDpackets: A Clean-label Backdoor Attack on Network Traffic Classifiers via Feature FusionMengxia Zhang, Yixiao Xu, Mohan Li, Yanbin Sun et al.INFOCOM 2026
- Realtime Robust Malicious Traffic Detection via Frequency Domain AnalysisChuanpu Fu, Qi Li, Meng Shen, Ke XuCCS 2021 · 194 citations
