Graph-based Incident Aggregation for Large-Scale Online Service Systems
Zhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang, Xuemin Wen, Xiao Ling, Yongqiang Yang, Michael R. Lyu
Abstract
As online service systems continue to grow in terms of complexity and volume, how service incidents are managed will significantly impact company revenue and user trust. Due to the cascading effect, cloud failures often come with an overwhelming number of incidents from dependent services and devices. To pursue efficient incident management, related incidents should be quickly aggregated to narrow down the problem scope. To this end, in this paper, we propose GRLIA, an incident aggregation framework based on graph representation learning over the cascading graph of cloud failures. A representation vector is learned for each unique type of incident in an unsupervised and unified manner, which is able to simultaneously encode the topological and temporal correlations among incidents. Thus, it can be easily employed for online incident aggregation. In particular, to learn the correlations more accurately, we try to recover the complete scope of failures’ cascading impact by leveraging fine-grained system monitoring data, i.e., Key Performance Indicators (KPIs). The proposed framework is evaluated with real-world incident data collected from a large-scale online service system of Huawei Cloud. The experimental results demonstrate that GRLIA is effective and outperforms existing methods. Furthermore, our framework has been successfully deployed in industrial practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43c11639-4c5a-4ee0-bdef-b75d9edddd38Cited by top-tier papers7
- Incident-aware Duplicate Ticket Aggregation for Cloud SystemsJinyang Liu, Shilin He, Zhuangbin Chen, Liqun Li et al.ICSE 2023 · 14 citations
- Prism: Revealing Hidden Functional Clusters from Massive Instances in Cloud SystemsJinyang Liu, Zhihan Jiang, Jiazhen Gu, Junjie Huang et al.ASE 2023 · 7 citations
- ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service SystemsWenwei Gu, Jiazhen Gu, Jinyang Liu, Zhuangbin Chen et al.ICSE 2025 · 4 citations
- Dynamic Graph Neural Networks-Based Alert Link Prediction for Online Service SystemsYiru Chen, Chenxi Zhang, Zhen Dong, Dingyu Yang et al.ASE 2023 · 3 citations
- Tracezip: Efficient Distributed Tracing via Trace CompressionZhuangbin Chen, Junsong Pu, Zibin ZhengISSTA 2025 · 2 citations
Builds on5
- Real-time incident prediction for online service systemsNengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng et al.FSE 2020 · 48 citations
- Identifying linked incidents in large-scale online service systemsYujun Chen, Xian Yang, Hang Dong, Xiaoting He et al.FSE 2020 · 43 citations
- Automatically and Adaptively Identifying Severe Alerts for Online Service SystemsNengwen Zhao, Panshi Jin, Lixin Wang, Xiaoqin Yang et al.INFOCOM 2020 · 36 citations
- Efficient incident identification from multi-dimensional issue reports via meta-heuristic searchJiazhen Gu, Chuan Luo, Si Qin, Bo Qiao et al.FSE 2020 · 35 citations
- Scouts: Improving the Diagnosis Process Through Domain-customized Incident RoutingJiaqi Gao, Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri et al.SIGCOMM 2020 · 18 citations
Related papers
- Graph based Incident Extraction and Diagnosis in Large-Scale Online SystemsZilong He, Pengfei Chen, Yu Luo, Qiuyu Yan et al.ASE 2022 · 12 citations
- Triangle: Empowering Incident Triage with Multi-AgentZhaoyang Yu, Aoyang Fang, Minghua Ma, Jaskaran Singh Walia et al.ASE 2025 · 2 citations
- AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud SystemsTianyi Yang, Jiacheng Shen, Yuxin Su, Xiao Ling et al.ASE 2021 · 23 citations
- Online Summarizing Alerts through Semantic and Behavior InformationJia Chen, Peng Wang, Wei WangICSE 2022 · 16 citations
- ART: A Unified Unsupervised Framework for Incident Management in Microservice SystemsYongqian Sun, Binpeng Shi, Mingyu Mao, Minghua Ma et al.ASE 2024 · 9 citations
