How Incidental are the Incidents? Characterizing and Prioritizing Incidents for Large-Scale Online Service Systems
Junjie Chen, Shu Zhang, Xiaoting He, Qingwei Lin, Hongyu Zhang, Dan Hao, Yu Kang, Feng Gao, Zhangwei Xu, Yingnong Dang, Dongmei Zhang
摘要
Although tremendous efforts have been devoted to the quality assurance of online service systems, in reality, these systems still come across many incidents (i.e., unplanned interruptions and outages), which can decrease user satisfaction or cause economic loss. To better understand the characteristics of incidents and improve the incident management process, we perform the first large-scale empirical analysis of incidents collected from 18 real-world online service systems in Microsoft. Surprisingly, we find that although a large number of incidents could occur over a short period of time, many of them actually do not matter, i.e., engineers will not fix them with a high priority after manually identifying their root cause. We call these incidents incidental incidents. Our qualitative and quantitative analyses show that incidental incidents are significant in terms of both number and cost. Therefore, it is important to prioritize incidents by identifying incidental incidents in advance to optimize incident management efforts. In particular, we propose an approach, called DeepIP (Deep learning based Incident Prioritization), to prioritizing incidents based on a large amount of historical incident data. More specifically, we design an attention-based Convolutional Neural Network (CNN) to learn a prediction model to identify incidental incidents. We then prioritize all incidents by ranking the predicted probabilities of incidents being incidental. We evaluate the performance of DeepIP using real-world incident data. The experimental results show that DeepIP effectively prioritizes incidents by identifying incidental incidents and significantly outperforms all the compared approaches. For example, the AUC of DeepIP achieves 0.808, while that of the best compared approach is only 0.624 on average.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Semi-supervised Log-based Anomaly Detection via Probabilistic Label EstimationLin Yang, Junjie Chen, Zan Wang, Weijing Wang 等ICSE 2021 · 被引用 216 次
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang 等EuroSys 2024 · 被引用 175 次
- Deep learning library testing via effective model generationZan Wang, Ming Yan, Junjie Chen, Shuang Liu 等FSE 2020 · 被引用 165 次
- Prioritizing Test Inputs for Deep Neural Networks via Mutation AnalysisZan Wang, Hanmo You, Junjie Chen, Yingyi Zhang 等ICSE 2021 · 被引用 117 次
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann 等ICSE 2023 · 被引用 93 次
它引用的顶会 Paper3
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 被引用 1,823 次
- Real-time incident prediction for online service systemsNengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng 等FSE 2020 · 被引用 48 次
- Identifying linked incidents in large-scale online service systemsYujun Chen, Xian Yang, Hang Dong, Xiaoting He 等FSE 2020 · 被引用 43 次
相关 Paper
- Efficient incident identification from multi-dimensional issue reports via meta-heuristic searchJiazhen Gu, Chuan Luo, Si Qin, Bo Qiao 等FSE 2020 · 被引用 35 次
- Fast Outage Analysis of Large-scale Production Clouds with Service Correlation MiningYaohui Wang, Guozheng Li, Zijian Wang, Yu Kang 等ICSE 2021 · 被引用 24 次
- Online Summarizing Alerts through Semantic and Behavior InformationJia Chen, Peng Wang, Wei WangICSE 2022 · 被引用 16 次
- Graph based Incident Extraction and Diagnosis in Large-Scale Online SystemsZilong He, Pengfei Chen, Yu Luo, Qiuyu Yan 等ASE 2022 · 被引用 12 次
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan 等ISSTA 2020 · 被引用 206 次
