Pluto: Sample Selection for Robust Anomaly Detection on Polluted Log Data
Lei Ma, Lei Cao, Peter M. VanNostrand, Dennis M. Hofmann, Yao Su, Elke A. Rundensteiner
摘要
Log anomaly detection, critical in identifying system failures and preempting security breaches, finds irregular patterns within large volumes of log data. Modern log anomaly detectors rely on training deep learning models on clean anomaly-free log data. However, such clean log data requires expensive and tedious human labeling. In this paper, we thus propose a robust log anomaly detection framework, PlutoNOSPACE, that automatically selects a clean representative sample subset of the polluted log sequence data to train a Transformer-based anomaly detection model. Pluto features three innovations. First, due to localized concentrations of anomalies inherent in the embedding space of log data, Pluto partitions the sequence embedding space generated by the model into regions that then allow it to identify and discard regions that are highly polluted by our pollution level estimation scheme, based on our pollution quantification via Gaussian mixture modeling. Second, for the remaining more slightly polluted regions, we select samples that maximally purify the eigenvector spectrum, which can be transformed into the NP-hard facility location problem; allowing us to leverage its greedy solution with a (1-(1/e)) approximation guarantee in optimality. Third, by iteratively alternating between the above subset selection, a model re-training on the latest subset, and a subset filtering using dynamic training artifacts generated by the latest model, the data selected is progressively refined. The final sample set is used to retrain the final anomaly detection model. Our experiments on four real-world log benchmark datasets demonstrate that by retaining 77.7% (BGL) to 96.6% (ThunderBird) of the normal sequences while effectively removing 90.3% (BGL) to 100.0% (ThunderBird, HDFS) of the anomalies, Pluto provides a significant absolute F-1 improvement up to 68.86% (2.16% → 71.02%) compared to the state-of-the-art sample selection methods. The implementation of this work is available at https://github.com/LeiMa0324/Pluto-SIGMOD25.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CoLA: Model Collaboration for Log-based Anomaly DetectionXuhang Zhu, Xiu Tang, Sai Wu, Jichen Li 等VLDB 2025 · 被引用 2 次
- Krone: Hierarchical and Modular Log Anomaly DetectionLei Ma, Jinyang Liu, Tieying Zhang, Peter M. VanNostrand 等ICDE 2026
它引用的顶会 Paper10
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 被引用 1,823 次
- Log-based Anomaly Detection with Deep Learning: How Far Are We?Van-Hoang Le, Hongyu ZhangICSE 2022 · 被引用 212 次
- FINE Samples for Learning with Noisy LabelsTaehyeon Kim, Jongwoo Ko, Sangwook Cho, Jinhwan Choi 等NeurIPS 2021 · 被引用 145 次
- A Topological Filter for Learning with Label NoisePengxiang Wu, Songzhu Zheng, Mayank Goswami, Dimitris N. Metaxas 等NeurIPS 2020 · 被引用 143 次
- Searching to Exploit Memorization Effect in Learning with Noisy LabelsQuanming Yao, Hansi Yang, Bo Han, Gang Niu 等ICML 2020 · 被引用 121 次
相关 Paper
- Semi-supervised Log-based Anomaly Detection via Probabilistic Label EstimationLin Yang, Junjie Chen, Zan Wang, Weijing Wang 等ICSE 2021 · 被引用 216 次
- MetaLog: Generalizable Cross-System Anomaly Detection from Logs with Meta-LearningChenyangguang Zhang, Tong Jia, Guopeng Shen, Pinyan Zhu 等ICSE 2024 · 被引用 28 次
- Log-based Anomaly Detection Without Log ParsingVan-Hoang Le, Hongyu ZhangASE 2021 · 被引用 249 次
- A Critical Review of Common Log Data Sets Used for Evaluation of Sequence-Based Anomaly Detection TechniquesMax Landauer, Florian Skopik, Markus WurzenbergerFSE 2024 · 被引用 44 次
- ADA: Adaptive Deep Log Anomaly DetectorYali Yuan, Sripriya Srikant Adhatarao, Mingkai Lin, Yachao Yuan 等INFOCOM 2020 · 被引用 58 次
