Automatically and Adaptively Identifying Severe Alerts for Online Service Systems
Nengwen Zhao, Panshi Jin, Lixin Wang, Xiaoqin Yang, Rong Liu, Wenchi Zhang, Kaixin Sui, Dan Pei
摘要
In large-scale online service system, to enhance the quality of services, engineers need to collect various monitoring data and write many rules to trigger alerts. However, the number of alerts is way more than what on-call engineers can properly investigate. Thus, in practice, alerts are classified into several priority levels using manual rules, and on-call engineers primarily focus on handling the alerts with the highest priority level (i.e., severe alerts). Unfortunately, due to the complex and dynamic nature of the online services, this rule-based approach results in missed severe alerts or wasted troubleshooting time on non-severe alerts. In this paper, we propose AlertRank, an automatic and adaptive framework for identifying severe alerts. Specifically, AlertRank extracts a set of powerful and interpretable features (textual and temporal alert features, univariate and multivariate anomaly features for monitoring metrics), adopts XGBoost ranking algorithm to identify the severe alerts out of all incoming alerts, and uses novel methods to obtain labels for both training and testing. Experiments on the datasets from a top global commercial bank demonstrate that AlertRank is effective and achieves the F1-score of 0.89 on average, outperforming all baselines. The feedback from practice shows AlertRank can significantly save the manual efforts for on-call engineers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- Identifying bad software changes via multimodal anomaly detection for online service systemsNengwen Zhao, Junjie Chen, Zhaoyang Yu, Honglin Wang 等FSE 2021 · 被引用 89 次
- Groot: An Event-graph-based Approach for Root Cause Analysis in Industrial SettingsHanzhang Wang, Zhengkai Wu, Huai Jiang, Yichao Huang 等ASE 2021 · 被引用 73 次
- Real-time incident prediction for online service systemsNengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng 等FSE 2020 · 被引用 48 次
- Graph-based Incident Aggregation for Large-Scale Online Service SystemsZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang 等ASE 2021 · 被引用 29 次
- Outage-Watch: Early Prediction of Outages using Extreme Event RegularizerShubham Agarwal, Sarthak Chakraborty, Shaddy Garg, Sumit Bisht 等FSE 2023 · 被引用 5 次
相关 Paper
- Dynamic Graph Neural Networks-Based Alert Link Prediction for Online Service SystemsYiru Chen, Chenxi Zhang, Zhen Dong, Dingyu Yang 等ASE 2023 · 被引用 3 次
- Online Summarizing Alerts through Semantic and Behavior InformationJia Chen, Peng Wang, Wei WangICSE 2022 · 被引用 16 次
- Adaptive Performance Anomaly Detection for Online Service Systems via Pattern SketchingZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang 等ICSE 2022 · 被引用 44 次
- Hybrid anomaly detection and prioritization for network logs at cloud scaleDavid Ohana, Bruno Wassermann, Nicolas Dupuis, Elliot K. Kolodner 等EuroSys 2022 · 被引用 6 次
- ESRO: Experience Assisted Service Reliability against OutagesSarthak Chakraborty, Shubham Agarwal, Shaddy Garg, Abhimanyu Sethia 等ASE 2023 · 被引用 3 次
