Hybrid anomaly detection and prioritization for network logs at cloud scale
David Ohana, Bruno Wassermann, Nicolas Dupuis, Elliot K. Kolodner, Eran Raichstein, Michal Malka
Abstract
Monitoring the health of large-scale systems requires significant manual effort, usually through the continuous curation of alerting rules based on keywords, thresholds and regular expressions, which might generate a flood of mostly irrelevant alerts and obscure the actual information operators would like to see. Existing approaches try to improve the observability of systems by intelligently detecting anomalous situations. Such solutions surface anomalies that are statistically significant, but may not represent events that reliability engineers consider relevant. We propose ADEPTUS, a practical approach for detection of relevant health issues in an established system. ADEPTUS combines statistics and unsupervised learning to detect anomalies with supervised learning and heuristics to determine which of the detected anomalies are likely to be relevant to the Site Reliability Engineers (SREs). ADEPTUS overcomes the labor-intensive prerequisite of obtaining anomaly labels for supervised learning by automatically extracting information from historic alerts and incident tickets. We leverage ADEPTUS for observability in the network infrastructure of IBM Cloud. We perform an extensive real-world evaluation on 10 months of logs generated by tens of thousands of network devices across 11 data centers and demonstrate that ADEPTUS achieves higher alerting accuracy than the rule-based log alerting solution, curated by domain experts, used by SREs daily.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 46c46817-d471-45ae-a2fc-53c154cbadb6Related papers
- Heterogeneous Anomaly Detection for Software Systems via Semi-supervised Cross-modal AttentionCheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su et al.ICSE 2023 · 52 citations
- ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service SystemsWenwei Gu, Jiazhen Gu, Jinyang Liu, Zhuangbin Chen et al.ICSE 2025 · 4 citations
- Semantic Curriculum for Anomaly Detection: A Unified Language-Driven Meta-Optimization FrameworkKai Tan, Yangliu Du, Dongyang Zhan, Haining Yu et al.INFOCOM 2026
- Online Summarizing Alerts through Semantic and Behavior InformationJia Chen, Peng Wang, Wei WangICSE 2022 · 16 citations
- Semi-supervised Log-based Anomaly Detection via Probabilistic Label EstimationLin Yang, Junjie Chen, Zan Wang, Weijing Wang et al.ICSE 2021 · 216 citations
