Actionable and interpretable fault localization for recurring failures in online service systems
Zeyan Li, Nengwen Zhao, Mingjie Li, Xianglin Lu, Lixin Wang, Dongdong Chang, Xiaohui Nie, Li Cao, Wenchi Zhang, Kaixin Sui, Yanhua Wang, Xu Du
摘要
Fault localization is challenging in an online service system due to its monitoring data's large volume and variety and complex dependencies across/within its components (e.g., services or databases). Furthermore, engineers require fault localization solutions to be actionable and interpretable, which existing research approaches cannot satisfy. Therefore, the common industry practice is that, for a specific online service system, its experienced engineers focus on localization for recurring failures based on the knowledge accumulated about the system and historical failures. More specifically, 1) they can identify the underlying root causes and take mitigation actions when pinpointing a group of indicative metrics on the faulty component; 2) their diagnosis knowledge is roughly based on how one failure might affect the components in the whole system.
Although the above common practice is actionable and interpretable, it is largely manual, thus slow and sometimes inaccurate. In this paper, we aim to automate this practice through machine learning. That is, we propose an actionable and interpretable fault localization approach, DéjàVu, for recurring failures in online service systems. For a specific online service system, DéjàVu takes historical failures and dependencies in the system as input and trains a localization model offline; for an incoming failure, the trained model online recommends where the failure occurs (i.e., the faulty components) and which kind of failure occurs (i.e., the indicative group of metrics) (thus actionable), which are further interpreted both globally and locally (thus interpretable). Based on the evaluation on 601 failures from three production systems and one open-source benchmark, in less than one second, DéjàVu can
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and ReasoningZhe Xie, Zeyan Li, Xiao He, Longlong Xu 等VLDB 2025 · 被引用 87 次
- D-Bot: Database Diagnosis System using Large Language ModelsXuanhe Zhou, Guoliang Li, Zhaoyan Sun, Zhiyuan Liu 等VLDB 2024 · 被引用 50 次
- Root Cause Analysis for Microservice System based on Causal Inference: How Far Are We?Luan Pham, Huong Ha, Hongyu ZhangASE 2024 · 被引用 14 次
- ART: A Unified Unsupervised Framework for Incident Management in Microservice SystemsYongqian Sun, Binpeng Shi, Mingyu Mao, Minghua Ma 等ASE 2024 · 被引用 9 次
- Giving Every Modality a Voice in Microservice Failure Diagnosis via Multimodal Adaptive OptimizationLei Tao, Shenglin Zhang, Zedong Jia, Jinrui Sun 等ASE 2024 · 被引用 7 次
它引用的顶会 Paper7
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 被引用 1,586 次
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo 等ASPLOS 2021 · 被引用 170 次
- MicroRank: End-to-End Latency Issue Localization with Extended Spectrum Analysis in Microservice EnvironmentsGuangba Yu, Pengfei Chen, Hongyang Chen, Zijie Guan 等WWW 2021 · 被引用 152 次
- AutoMAP: Diagnose Your Microservice-based Web Applications AutomaticallyMeng Ma, Jingmin Xu, Yuan Wang, Pengfei Chen 等WWW 2020 · 被引用 144 次
- Diagnosing Root Causes of Intermittent Slow Queries in Large-Scale Cloud DatabasesMinghua Ma, Zheng Yin, Shenglin Zhang, Sheng Wang 等VLDB 2020 · 被引用 119 次
相关 Paper
- Dynamic Graph Neural Networks-Based Alert Link Prediction for Online Service SystemsYiru Chen, Chenxi Zhang, Zhen Dong, Dingyu Yang 等ASE 2023 · 被引用 3 次
- SLIM: a Scalable and Interpretable Light-weight Fault Localization Algorithm for Imbalanced Data in MicroserviceRui Ren, Jingbang Yang, Linxiao Yang, Xinyue Gu 等ASE 2024 · 被引用 1 次
- Real-time incident prediction for online service systemsNengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng 等FSE 2020 · 被引用 48 次
- FaultInsight: Interpreting Hyperscale Data Center Host FaultsTingzhu Bi, Yang Zhang, Yicheng Pan, Yu Zhang 等KDD 2024 · 被引用 1 次
- Nezha: Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability DataGuangba Yu, Pengfei Chen, Yufeng Li, Hongyang Chen 等FSE 2023 · 被引用 131 次
