Faster, deeper, easier: crowdsourcing diagnosis of microservice kernel failure from user space
Yicheng Pan, Meng Ma, Xinrui Jiang, Ping Wang
Abstract
With the widespread use of cloud-native architecture, increasing web applications (apps) choose to build on microservices. Simultaneously, troubleshooting becomes full of challenges owing to the high dynamics and complexity of anomaly propagation. Existing diagnostic methods rely heavily on monitoring metrics collected from the kernel side of microservice systems. Without a comprehensive monitoring infrastructure, application owners and even cloud operators cannot resort to these kernel-space solutions. This paper summarizes several insights on operating a top commercial cloud platform. Then, for the first time, we put forward the idea of user-space diagnosis for microservice kernel failures. To this end, we develop a crowdsourcing solution - DyCause, to resolve the asymmetric diagnostic information problem. DyCause deploys on the application side in a distributed manner. Through lightweight API log sharing, apps collect the operational status of kernel services collaboratively and initiate diagnosis on demand. Deploying DyCause is fast and lightweight as we do not have any architectural and functional requirements for the kernel. To reveal more accurate correlations from asymmetric diagnostic information, we design a novel statistical algorithm that can efficiently discover the time-varying causalities between services. This algorithm also helps us build the temporal order of the anomaly propagation. Therefore, by using DyCause, we can obtain more in-depth and interpretable diagnostic clues with limited indicators. We apply and evaluate DyCause on both a simulated test-bed and a real-world cloud system. Experimental results verify that DyCause running in the user-space outperforms several state-of-the-art algorithms running in the kernel on accuracy. Besides, DyCause shows superior advantages in terms of algorithmic efficiency and data sensitivity. Simply put, DyCause produces a significantly better result than other baselines when analyzing much fewer or sparser metrics. To conclude, DyCause is faster to act, deeper in analysis, and easier to deploy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 60d6229b-0e9d-472f-8ee0-2fa4a4a3ad94Cited by top-tier papers4
- DeepTraLog: Trace-Log Combined Microservice Anomaly Detection through Graph-based Deep LearningChenxi Zhang, Xin Peng, Chaofeng Sha, Ke Zhang et al.ICSE 2022 · 163 citations
- BARO: Robust Root Cause Analysis for Microservices via Multivariate Bayesian Online Change Point DetectionLuan Pham, Huong Ha, Hongyu ZhangFSE 2024 · 21 citations
- Root Cause Analysis for Microservice System based on Causal Inference: How Far Are We?Luan Pham, Huong Ha, Hongyu ZhangASE 2024 · 14 citations
- Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware BenchmarkAoyang Fang, Songhan Zhang, Yifan Yang, Haotong Wu et al.FSE 2026 · 1 citation
Related papers
- AutoMAP: Diagnose Your Microservice-based Web Applications AutomaticallyMeng Ma, Jingmin Xu, Yuan Wang, Pengfei Chen et al.WWW 2020 · 144 citations
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini et al.NeurIPS 2022 · 185 citations
- Look Deep into the Microservice System Anomaly through Very Sparse LogsXinrui Jiang, Yicheng Pan, Meng Ma, Ping WangWWW 2023 · 16 citations
- MRCA: Metric-level Root Cause Analysis for Microservices via Multi-Modal DataYidan Wang, Zhouruixing Zhu, Qiuai Fu, Yuchi Ma et al.ASE 2024 · 6 citations
- MicroRank: End-to-End Latency Issue Localization with Extended Spectrum Analysis in Microservice EnvironmentsGuangba Yu, Pengfei Chen, Hongyang Chen, Zijie Guan et al.WWW 2021 · 152 citations
