FARE: Enabling Fine-grained Attack Categorization under Low-quality Labeled Data
Junjie Liang, Wenbo Guo, Tongbo Luo, Vasant G. Honavar, Gang Wang, Xinyu Xing
摘要
—Supervised machine learning classifiers have been widely used for attack detection, but their training requires abundant high-quality labels. Unfortunately, high-quality labels are difficult to obtain in practice due to the high cost of data labeling and the constant evolution of attackers. Without such labels, it is challenging to train and deploy targeted countermeasures. In this paper, we propose FARE , a clustering method to enable fine-grained attack categorization under low-quality labels. We focus on two common issues in data labels: 1) missing labels for certain attack classes or families; and 2) only having coarse-grained labels available for different attack types. The core idea of FARE is to take full advantage of the limited labels while using the underlying data distribution to consolidate the low-quality labels. We design an ensemble model to fuse the results of multiple unsupervised learning algorithms with the given labels to mitigate the negative impact of missing classes and coarse-grained labels. We then train an input transformation network to map the input data into a low-dimensional latent space for fine-grained clustering. Using two security datasets (Android malware and network intrusion traces), we show that FARE significantly outperforms the state-of-the-art (semi-)supervised learning methods in clustering quality/correctness. Further, we perform an initial deployment of FARE by working with a large e-commerce service to detect fraudulent accounts. With real-world A/B tests and manual investigation, we demonstrate the effectiveness of FARE to catch previously-unseen frauds.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DeepAID: Interpreting and Improving Deep Learning-based Anomaly Detection in Security ApplicationsDongqi Han, Zhiliang Wang, Wenqi Chen, Ying Zhong 等CCS 2021 · 被引用 108 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- Learning from Limited Heterogeneous Training Data: Meta-Learning for Unsupervised Zero-Day Web Attack Detection across Web DomainsPeiyang Li, Ye Wang, Qi Li, Zhuotao Liu 等CCS 2023 · 被引用 13 次
- Dos and Don'ts of Machine Learning in Computer SecurityDaniel Arp, Erwin Quiring, Feargus Pendlebury, Alexander Warnecke 等USENIX Security 2022
- Revisiting Concept Drift in Windows Malware Detection: Adaptation to Real Drifted Malware with Minimal SamplesAdrian Shuai Li, Arun Iyengar, Ashish Kundu, Elisa BertinoNDSS 2025
它引用的顶会 Paper7
- Kitsune: An Ensemble of Autoencoders for Online Network Intrusion DetectionYisroel Mirsky, Tomer Doitshman, Yuval Elovici, Asaf ShabtaiNDSS 2018 · 被引用 945 次
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder 等USENIX Security 2019 · 被引用 441 次
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang 等USENIX Security 2017 · 被引用 325 次
- Practical Attacks Against Graph-based ClusteringYizheng Chen, Yacin Nadji, Athanasios Kountouras, Fabian Monrose 等CCS 2017 · 被引用 90 次
- Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data AugmentationSteve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu 等S&P 2020 · 被引用 88 次
相关 Paper
- Fine-grained Classes and How to Find ThemMatej Grcic, Artyom Gadetsky, Maria BrbicICML 2024 · 被引用 5 次
- Adaptive Clustering-based Malicious Traffic Classification at the Network EdgeAlec F. Diallo, Paul PatrasINFOCOM 2021 · 被引用 64 次
- Training with Only 1.0 ‰ Samples: Malicious Traffic Detection via Cross-Modality Feature FusionChuanpu Fu, Qi Li, Elisa Bertino, Ke XuCCS 2025
- Towards Cross-Granularity Few-Shot Learning: Coarse-to-Fine Pseudo-Labeling with Visual-Semantic Meta-EmbeddingJinhai Yang, Hua Yang, Lin ChenACM MM 2021 · 被引用 16 次
- A Generic Method for Fine-grained Category Discovery in Natural Language TextsChang Tian, Matthew B. Blaschko, Wenpeng Yin, Mingzhe Xing 等EMNLP 2024 · 被引用 2 次
