Real-time incident prediction for online service systems
Nengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng, Gang Wang, Yong Wu, Fang Zhou, Zhen Feng, Xiaohui Nie, Wenchi Zhang, Kaixin Sui, Dan Pei
摘要
Incidents in online service systems could dramatically degrade system availability and destroy user experience. To guarantee service quality and reduce economic loss, it is essential to predict the occurrence of incidents in advance so that engineers can take some proactive actions to prevent them. In this work, we propose an effective and interpretable incident prediction approach, called eWarn, which utilizes historical data to forecast whether an incident will happen in the near future based on alert data in real time. More specifically, eWarn first extracts a set of effective features (including textual features and statistical features) to represent omen alert patterns via careful feature engineering. To reduce the influence of noisy alerts (that are not relevant to the occurrence of incidents), eWarn then incorporates the multi-instance learning formulation. Finally, eWarn builds a classification model via machine learning and generates an interpretable report about the prediction result via a state-of-the-art explanation technique (i.e., LIME). In this way, an early warning signal along with its interpretable report can be sent to engineers to facilitate their understanding and handling for the incoming incident. An extensive study on 11 real-world online service systems from a large commercial bank demonstrates the effectiveness of eWarn, outperforming state-of-the-art alert-based incident prediction approaches and the practice of incident prediction with alerts. In particular, we have applied eWarn to two large commercial banks in practice and shared some success stories and lessons learned from real deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Semi-supervised Log-based Anomaly Detection via Probabilistic Label EstimationLin Yang, Junjie Chen, Zan Wang, Weijing Wang 等ICSE 2021 · 被引用 216 次
- Identifying bad software changes via multimodal anomaly detection for online service systemsNengwen Zhao, Junjie Chen, Zhaoyang Yu, Honglin Wang 等FSE 2021 · 被引用 89 次
- Fighting the Fog of War: Automated Incident Detection for Cloud SystemsLiqun Li, Xu Zhang, Xin Zhao, Hongyu Zhang 等USENIX ATC 2021 · 被引用 43 次
- Recommending Good First Issues in GitHub OSS ProjectsWenxin Xiao, Hao He, Weiwei Xu, Xin Tan 等ICSE 2022 · 被引用 35 次
- How Incidental are the Incidents? Characterizing and Prioritizing Incidents for Large-Scale Online Service SystemsJunjie Chen, Shu Zhang, Xiaoting He, Qingwei Lin 等ASE 2020 · 被引用 33 次
它引用的顶会 Paper2
- Automatically and Adaptively Identifying Severe Alerts for Online Service SystemsNengwen Zhao, Panshi Jin, Lixin Wang, Xiaoqin Yang 等INFOCOM 2020 · 被引用 36 次
- How Incidental are the Incidents? Characterizing and Prioritizing Incidents for Large-Scale Online Service SystemsJunjie Chen, Shu Zhang, Xiaoting He, Qingwei Lin 等ASE 2020 · 被引用 33 次
相关 Paper
- Identifying linked incidents in large-scale online service systemsYujun Chen, Xian Yang, Hang Dong, Xiaoting He 等FSE 2020 · 被引用 43 次
- Online Summarizing Alerts through Semantic and Behavior InformationJia Chen, Peng Wang, Wei WangICSE 2022 · 被引用 16 次
- Alert Summarization for Online Service Systems by Validating Propagation Paths of FaultsJia Chen, Yuang He, Peng Wang, Xiaolei Chen 等FSE 2025
- Actionable and interpretable fault localization for recurring failures in online service systemsZeyan Li, Nengwen Zhao, Mingjie Li, Xianglin Lu 等FSE 2022 · 被引用 69 次
- Imminence Monitoring of Critical Events: A Representation Learning ApproachYan Li, Tingjian GeSIGMOD 2021 · 被引用 6 次
