USENIX Security2023Top-tier venue
Generative Intrusion Detection and Prevention on Data Stream
HyungBin Seo, MyungKeun Yoon
Abstract
Data arrive in a stream, for example, network packets, emails, or malicious files, and ideally they should be investigated for cybersecurity. The current best practice would be to check if each data includes any suspicious signatures, or simply strings, which were obtained a priori by elaborate manual analysis in previous cyberattack cases. Unfortunately, unknown attacks, called zero-day attacks, cannot be timely detected in this way because no signature is available yet. To tackle this problem, recent studies have presented high-speed methods that can extract frequent substrings from the data stream and use them as attack signatures because the frequently-occurred signatures are often related with attacks; unfortunately, more benign signatures are extracted than malicious ones, especially when there is no attack in most of the time. This causes both a tremendous number of false-positives and extra human interventions to remove benign signatures. In this paper, we design a new streaming algorithm that can first identify a frequent group of signatures appearing together at the same time from data streams. Using this frequent signature-group instead of frequently-occurred individual signatures, the new scheme achieves a high detection accuracy by mitigating the false-positive problem with only a small fixed amount of memory and a constant number of hash operations, which has not been achieved by any previous work. This improvement comes from a new method for summarizing similar data with a fixed amount of memory, called a minHashed virtual vector, which allows us to automatically identify a frequent group of signatures with each data read only once. We perform exhaustive experiments on different private and open datasets, to verify both the practical effectiveness and the experimental reproducibility of the new scheme.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5cc7be9-80a6-402a-8b24-08f288a254b3Cited by top-tier papers1
Ask how each one uses itBuilds on8
- NoDoze: Combatting Threat Alert Fatigue with Automated Provenance TriageWajih Ul Hassan, Shengjian Guo, Ding Li, Zhengzhang Chen et al.NDSS 2019 · 411 citations
- Matched and Mismatched SOCs: A Qualitative Study on Security Operations Center IssuesFaris Bugra Kokulu, Ananta Soneji, Tiffany Bao, Yan Shoshitaishvili et al.CCS 2019 · 134 citations
- ZeroWall: Detecting Zero-Day Web Attacks through Encoder-Decoder Recurrent Neural NetworksRuming Tang, Zheng Yang, Zeyan Li, Weibin Meng et al.INFOCOM 2020 · 97 citations
- DEEPCASE: Semi-Supervised Contextual Analysis of Security EventsThijs van Ede, Hojjat Aghakhani, Noah Spahn, Riccardo Bortolameotti et al.S&P 2022 · 91 citations
- Adaptive Clustering-based Malicious Traffic Classification at the Network EdgeAlec F. Diallo, Paul PatrasINFOCOM 2021 · 64 citations
Related papers
- Scout Sketch: Finding Promising Items in Data StreamsTianyu Ma, Guoju Gao, He Huang, Yu-e Sun et al.INFOCOM 2024 · 4 citations
- Sketch-Based Anomaly Detection in Streaming GraphsSiddharth Bhatia, Mohit Wadhwa, Kenji Kawaguchi, Neil Shah et al.KDD 2023 · 23 citations
- Finding Simplex Items in Data StreamsZhuochen Fan, Jiarui Guo, Xiaodong Li, Tong Yang et al.ICDE 2023 · 9 citations
- Midas: Microcluster-Based Detector of Anomalies in Edge StreamsSiddharth Bhatia, Bryan Hooi, Minji Yoon, Kijung Shin et al.AAAI 2020 · 118 citations
- CND-IDS: Continual Novelty Detection for Intrusion Detection SystemsSean Fuhrman, Onat Güngör, Tajana RosingDAC 2025 · 9 citations
