Graph based Incident Extraction and Diagnosis in Large-Scale Online Systems
Zilong He, Pengfei Chen, Yu Luo, Qiuyu Yan, Hongyang Chen, Guangba Yu, Fangyuan Li
Abstract
With the ever increasing scale and complexity of online systems, incidents are gradually becoming commonplace. Without appropriate handling, they can seriously harm the system availability. However, in large-scale online systems, these incidents are usually drowning in a slew of issues (i.e., something abnormal, while not necessarily an incident), rendering them difficult to handle. Typically, these issues will result in a cascading effect across the system, and a proper management of the incidents depends heavily on a thorough analysis of this effect. Therefore, in this paper, we propose a method to automatically analyze the cascading effect of availability issues in online systems and extract the corresponding graph based issue representations incorporating both of the issue symptoms and affected service attributes. With the extracted representations, we train and utilize a graph neural networks based model to perform incident detection. Then, for the detected incident, we leverage the PageRank algorithm with a flexible transition matrix design to locate its root cause. We evaluate our approach using real-world data collected from the WeChat online service system, the largest instant message system in China. The results confirm the effectiveness of our approach. Moreover, our approach is successfully deployed in the company and eases the burden of operators in the face of a flood of issues and related alert signals.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 35a05b70-22a1-4b69-97df-52dd54da852eCited by top-tier papers9
- Xpert: Empowering Incident Management with Query Recommendations via Large Language ModelsYuxuan Jiang, Chaoyun Zhang, Shilin He, Zhihao Yang et al.ICSE 2024 · 24 citations
- BARO: Robust Root Cause Analysis for Microservices via Multivariate Bayesian Online Change Point DetectionLuan Pham, Huong Ha, Hongyu ZhangFSE 2024 · 21 citations
- Root Cause Analysis for Microservice System based on Causal Inference: How Far Are We?Luan Pham, Huong Ha, Hongyu ZhangASE 2024 · 14 citations
- ESRO: Experience Assisted Service Reliability against OutagesSarthak Chakraborty, Shubham Agarwal, Shaddy Garg, Abhimanyu Sethia et al.ASE 2023 · 3 citations
- TORAI: Multi-source Root Cause Analysis for Blind Spots in Microservice Service Call GraphLuan Pham, Huong Ha, Xiuzhen Zhang, Hongyu ZhangFSE 2026 · 3 citations
Related papers
- ChangeRCA: Finding Root Causes from Software Changes in Large Online SystemsGuangba Yu, Pengfei Chen, Zilong He, Qiuyu Yan et al.FSE 2024 · 3 citations
- Dynamic Graph Neural Networks-Based Alert Link Prediction for Online Service SystemsYiru Chen, Chenxi Zhang, Zhen Dong, Dingyu Yang et al.ASE 2023 · 3 citations
- Graph-based Incident Aggregation for Large-Scale Online Service SystemsZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang et al.ASE 2021 · 29 citations
- Identifying linked incidents in large-scale online service systemsYujun Chen, Xian Yang, Hang Dong, Xiaoting He et al.FSE 2020 · 43 citations
- How Incidental are the Incidents? Characterizing and Prioritizing Incidents for Large-Scale Online Service SystemsJunjie Chen, Shu Zhang, Xiaoting He, Qingwei Lin et al.ASE 2020 · 33 citations
