Actionable and interpretable fault localization for recurring failures in online service systems
Zeyan Li, Nengwen Zhao, Mingjie Li, Xianglin Lu, Lixin Wang, Dongdong Chang, Xiaohui Nie, Li Cao, Wenchi Zhang, Kaixin Sui, Yanhua Wang, Xu Du
Abstract
Fault localization is challenging in an online service system due to its monitoring data's large volume and variety and complex dependencies across/within its components (e.g., services or databases). Furthermore, engineers require fault localization solutions to be actionable and interpretable, which existing research approaches cannot satisfy. Therefore, the common industry practice is that, for a specific online service system, its experienced engineers focus on localization for recurring failures based on the knowledge accumulated about the system and historical failures. More specifically, 1) they can identify the underlying root causes and take mitigation actions when pinpointing a group of indicative metrics on the faulty component; 2) their diagnosis knowledge is roughly based on how one failure might affect the components in the whole system.
Although the above common practice is actionable and interpretable, it is largely manual, thus slow and sometimes inaccurate. In this paper, we aim to automate this practice through machine learning. That is, we propose an actionable and interpretable fault localization approach, DéjàVu, for recurring failures in online service systems. For a specific online service system, DéjàVu takes historical failures and dependencies in the system as input and trains a localization model offline; for an incoming failure, the trained model online recommends where the failure occurs (i.e., the faulty components) and which kind of failure occurs (i.e., the indicative group of metrics) (thus actionable), which are further interpreted both globally and locally (thus interpretable). Based on the evaluation on 601 failures from three production systems and one open-source benchmark, in less than one second, DéjàVu can
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2127521-da15-4c6b-8322-7d1fa16b75ccCited by top-tier papers12
- ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and ReasoningZhe Xie, Zeyan Li, Xiao He, Longlong Xu et al.VLDB 2025 · 87 citations
- D-Bot: Database Diagnosis System using Large Language ModelsXuanhe Zhou, Guoliang Li, Zhaoyan Sun, Zhiyuan Liu et al.VLDB 2024 · 50 citations
- Root Cause Analysis for Microservice System based on Causal Inference: How Far Are We?Luan Pham, Huong Ha, Hongyu ZhangASE 2024 · 14 citations
- ART: A Unified Unsupervised Framework for Incident Management in Microservice SystemsYongqian Sun, Binpeng Shi, Mingyu Mao, Minghua Ma et al.ASE 2024 · 9 citations
- Giving Every Modality a Voice in Microservice Failure Diagnosis via Multimodal Adaptive OptimizationLei Tao, Shenglin Zhang, Zedong Jia, Jinrui Sun et al.ASE 2024 · 7 citations
Builds on7
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo et al.ASPLOS 2021 · 170 citations
- MicroRank: End-to-End Latency Issue Localization with Extended Spectrum Analysis in Microservice EnvironmentsGuangba Yu, Pengfei Chen, Hongyang Chen, Zijie Guan et al.WWW 2021 · 152 citations
- AutoMAP: Diagnose Your Microservice-based Web Applications AutomaticallyMeng Ma, Jingmin Xu, Yuan Wang, Pengfei Chen et al.WWW 2020 · 144 citations
- Diagnosing Root Causes of Intermittent Slow Queries in Large-Scale Cloud DatabasesMinghua Ma, Zheng Yin, Shenglin Zhang, Sheng Wang et al.VLDB 2020 · 119 citations
Related papers
- Dynamic Graph Neural Networks-Based Alert Link Prediction for Online Service SystemsYiru Chen, Chenxi Zhang, Zhen Dong, Dingyu Yang et al.ASE 2023 · 3 citations
- SLIM: a Scalable and Interpretable Light-weight Fault Localization Algorithm for Imbalanced Data in MicroserviceRui Ren, Jingbang Yang, Linxiao Yang, Xinyue Gu et al.ASE 2024 · 1 citation
- Real-time incident prediction for online service systemsNengwen Zhao, Junjie Chen, Zhou Wang, Xiao Peng et al.FSE 2020 · 48 citations
- FaultInsight: Interpreting Hyperscale Data Center Host FaultsTingzhu Bi, Yang Zhang, Yicheng Pan, Yu Zhang et al.KDD 2024 · 1 citation
- Nezha: Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability DataGuangba Yu, Pengfei Chen, Yufeng Li, Hongyang Chen et al.FSE 2023 · 131 citations
