Hybrid anomaly detection and prioritization for network logs at cloud scale
David Ohana, Bruno Wassermann, Nicolas Dupuis, Elliot K. Kolodner, Eran Raichstein, Michal Malka
摘要
Monitoring the health of large-scale systems requires significant manual effort, usually through the continuous curation of alerting rules based on keywords, thresholds and regular expressions, which might generate a flood of mostly irrelevant alerts and obscure the actual information operators would like to see. Existing approaches try to improve the observability of systems by intelligently detecting anomalous situations. Such solutions surface anomalies that are statistically significant, but may not represent events that reliability engineers consider relevant. We propose ADEPTUS, a practical approach for detection of relevant health issues in an established system. ADEPTUS combines statistics and unsupervised learning to detect anomalies with supervised learning and heuristics to determine which of the detected anomalies are likely to be relevant to the Site Reliability Engineers (SREs). ADEPTUS overcomes the labor-intensive prerequisite of obtaining anomaly labels for supervised learning by automatically extracting information from historic alerts and incident tickets. We leverage ADEPTUS for observability in the network infrastructure of IBM Cloud. We perform an extensive real-world evaluation on 10 months of logs generated by tens of thousands of network devices across 11 data centers and demonstrate that ADEPTUS achieves higher alerting accuracy than the rule-based log alerting solution, curated by domain experts, used by SREs daily.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Heterogeneous Anomaly Detection for Software Systems via Semi-supervised Cross-modal AttentionCheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su 等ICSE 2023 · 被引用 52 次
- ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service SystemsWenwei Gu, Jiazhen Gu, Jinyang Liu, Zhuangbin Chen 等ICSE 2025 · 被引用 4 次
- Semantic Curriculum for Anomaly Detection: A Unified Language-Driven Meta-Optimization FrameworkKai Tan, Yangliu Du, Dongyang Zhan, Haining Yu 等INFOCOM 2026
- Online Summarizing Alerts through Semantic and Behavior InformationJia Chen, Peng Wang, Wei WangICSE 2022 · 被引用 16 次
- Semi-supervised Log-based Anomaly Detection via Probabilistic Label EstimationLin Yang, Junjie Chen, Zan Wang, Weijing Wang 等ICSE 2021 · 被引用 216 次
