FARE: Enabling Fine-grained Attack Categorization under Low-quality Labeled Data
Junjie Liang, Wenbo Guo, Tongbo Luo, Vasant G. Honavar, Gang Wang, Xinyu Xing
Abstract
—Supervised machine learning classifiers have been widely used for attack detection, but their training requires abundant high-quality labels. Unfortunately, high-quality labels are difficult to obtain in practice due to the high cost of data labeling and the constant evolution of attackers. Without such labels, it is challenging to train and deploy targeted countermeasures. In this paper, we propose FARE , a clustering method to enable fine-grained attack categorization under low-quality labels. We focus on two common issues in data labels: 1) missing labels for certain attack classes or families; and 2) only having coarse-grained labels available for different attack types. The core idea of FARE is to take full advantage of the limited labels while using the underlying data distribution to consolidate the low-quality labels. We design an ensemble model to fuse the results of multiple unsupervised learning algorithms with the given labels to mitigate the negative impact of missing classes and coarse-grained labels. We then train an input transformation network to map the input data into a low-dimensional latent space for fine-grained clustering. Using two security datasets (Android malware and network intrusion traces), we show that FARE significantly outperforms the state-of-the-art (semi-)supervised learning methods in clustering quality/correctness. Further, we perform an initial deployment of FARE by working with a large e-commerce service to detect fraudulent accounts. With real-world A/B tests and manual investigation, we demonstrate the effectiveness of FARE to catch previously-unseen frauds.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security ApplicationsDongqi Han, Zhiliang Wang, Wenqi Chen, Ying Zhong et al.CCS 2021 · 108 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Learning from Limited Heterogeneous Training Data: Meta-Learning for Unsupervised Zero-Day Web Attack Detection across Web DomainsPeiyang Li, Ye Wang, Qi Li, Zhuotao Liu et al.CCS 2023 · 13 citations
- Dos and Don'ts of Machine Learning in Computer SecurityDaniel Arp, Erwin Quiring, Feargus Pendlebury, Alexander Warnecke et al.USENIX Security 2022
- Revisiting Concept Drift in Windows Malware Detection: Adaptation to Real Drifted Malware with Minimal SamplesAdrian Shuai Li, Arun Iyengar, Ashish Kundu, Elisa BertinoNDSS 2025
Builds on7
- Kitsune: An Ensemble of Autoencoders for Online Network Intrusion DetectionYisroel Mirsky, Tomer Doitshman, Yuval Elovici, Asaf ShabtaiNDSS 2018 · 945 citations
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder et al.USENIX Security 2019 · 441 citations
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang et al.USENIX Security 2017 · 325 citations
- Practical Attacks Against Graph-based ClusteringYizheng Chen, Yacin Nadji, Athanasios Kountouras, Fabian Monrose et al.CCS 2017 · 90 citations
- Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data AugmentationSteve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu et al.S&P 2020 · 88 citations
Related papers
- Fine-grained Classes and How to Find ThemMatej Grcic, Artyom Gadetsky, Maria BrbicICML 2024 · 5 citations
- Adaptive Clustering-based Malicious Traffic Classification at the Network EdgeAlec F. Diallo, Paul PatrasINFOCOM 2021 · 64 citations
- Training with Only 1.0 ‰ Samples: Malicious Traffic Detection via Cross-Modality Feature FusionChuanpu Fu, Qi Li, Elisa Bertino, Ke XuCCS 2025
- Towards Cross-Granularity Few-Shot Learning: Coarse-to-Fine Pseudo-Labeling with Visual-Semantic Meta-EmbeddingJinhai Yang, Hua Yang, Lin ChenACM MM 2021 · 16 citations
- A Generic Method for Fine-grained Category Discovery in Natural Language TextsChang Tian, Matthew B. Blaschko, Wenpeng Yin, Mingzhe Xing et al.EMNLP 2024 · 2 citations
