Graph-based Incident Aggregation for Large-Scale Online Service Systems
Zhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang, Xuemin Wen, Xiao Ling, Yongqiang Yang, Michael R. Lyu
摘要
As online service systems continue to grow in terms of complexity and volume, how service incidents are managed will significantly impact company revenue and user trust. Due to the cascading effect, cloud failures often come with an overwhelming number of incidents from dependent services and devices. To pursue efficient incident management, related incidents should be quickly aggregated to narrow down the problem scope. To this end, in this paper, we propose GRLIA, an incident aggregation framework based on graph representation learning over the cascading graph of cloud failures. A representation vector is learned for each unique type of incident in an unsupervised and unified manner, which is able to simultaneously encode the topological and temporal correlations among incidents. Thus, it can be easily employed for online incident aggregation. In particular, to learn the correlations more accurately, we try to recover the complete scope of failures’ cascading impact by leveraging fine-grained system monitoring data, i.e., Key Performance Indicators (KPIs). The proposed framework is evaluated with real-world incident data collected from a large-scale online service system of Huawei Cloud. The experimental results demonstrate that GRLIA is effective and outperforms existing methods. Furthermore, our framework has been successfully deployed in industrial practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Incident-aware Duplicate Ticket Aggregation for Cloud SystemsJinyang Liu, Shilin He, Zhuangbin Chen, Liqun Li 等ICSE 2023 · 被引用 14 次
- Prism: Revealing Hidden Functional Clusters from Massive Instances in Cloud SystemsJinyang Liu, Zhihan Jiang, Jiazhen Gu, Junjie Huang 等ASE 2023 · 被引用 7 次
- ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service SystemsWenwei Gu, Jiazhen Gu, Jinyang Liu, Zhuangbin Chen 等ICSE 2025 · 被引用 4 次
- Dynamic Graph Neural Networks-Based Alert Link Prediction for Online Service SystemsYiru Chen, Chenxi Zhang, Zhen Dong, Dingyu Yang 等ASE 2023 · 被引用 3 次
- Tracezip: Efficient Distributed Tracing via Trace CompressionZhuangbin Chen, Junsong Pu, Zibin ZhengISSTA 2025 · 被引用 2 次
它引用的顶会 Paper5
- Real-time incident prediction for online service systemsNengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng 等FSE 2020 · 被引用 48 次
- Identifying linked incidents in large-scale online service systemsYujun Chen, Xian Yang, Hang Dong, Xiaoting He 等FSE 2020 · 被引用 43 次
- Automatically and Adaptively Identifying Severe Alerts for Online Service SystemsNengwen Zhao, Panshi Jin, Lixin Wang, Xiaoqin Yang 等INFOCOM 2020 · 被引用 36 次
- Efficient incident identification from multi-dimensional issue reports via meta-heuristic searchJiazhen Gu, Chuan Luo, Si Qin, Bo Qiao 等FSE 2020 · 被引用 35 次
- Scouts: Improving the Diagnosis Process Through Domain-customized Incident RoutingJiaqi Gao, Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri 等SIGCOMM 2020 · 被引用 18 次
相关 Paper
- Graph based Incident Extraction and Diagnosis in Large-Scale Online SystemsZilong He, Pengfei Chen, Yu Luo, Qiuyu Yan 等ASE 2022 · 被引用 12 次
- Triangle: Empowering Incident Triage with Multi-AgentZhaoyang Yu, Aoyang Fang, Minghua Ma, Jaskaran Singh Walia 等ASE 2025 · 被引用 2 次
- AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud SystemsTianyi Yang, Jiacheng Shen, Yuxin Su, Xiao Ling 等ASE 2021 · 被引用 23 次
- Online Summarizing Alerts through Semantic and Behavior InformationJia Chen, Peng Wang, Wei WangICSE 2022 · 被引用 16 次
- ART: A Unified Unsupervised Framework for Incident Management in Microservice SystemsYongqian Sun, Binpeng Shi, Mingyu Mao, Minghua Ma 等ASE 2024 · 被引用 9 次
