Murphy: Performance Diagnosis of Distributed Cloud Applications
Vipul Harsh, Wenxuan Zhou, Sachin Ashok, Radhika Niranjan Mysore, Brighten Godfrey, Sujata Banerjee
摘要
Modern cloud-based applications have complex inter-dependencies on both distributed application components as well as network infrastructure, making it difficult to reason about their performance. As a result, a rich body of work seeks to automate performance diagnosis of enterprise networks and such cloud applications. However, existing methods either ignore inter-dependencies which results in poor accuracy, or require causal acyclic dependencies which cannot model common enterprise environments.
We describe the design and implementation of Murphy, an automated performance diagnosis system, that can work with commonly available telemetry in practical enterprise environments, while achieving high accuracy. Murphy utilizes loosely-defined associations between entities obtained from commonly available monitoring data. Its learning algorithm is based on a Markov Random Field (MRF) that can take advantage of such loose associations to reason about how entities affect each other in the context of a specific incident. We evaluate Murphy in an emulated microservice environment and in real incidents from a large enterprise. Compared to past work, Murphy is able to reduce diagnosis error by ≈ 1.35× in restrictive environments supported by past work, and by ≥ 4.7× in more general environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- GAMMA: Graph Neural Network-Based Multi-Bottleneck Localization for Microservices ApplicationsGagan Somashekar, Anurag Dutt, Mainak Adak, Tania Lorido-Botran 等WWW 2024 · 被引用 15 次
- Towards LLM-Based Failure Localization in Production-Scale NetworksChenxu Wang, Xumiao Zhang, Runwei Lu, Xianshang Lin 等SIGCOMM 2025 · 被引用 13 次
- SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model TrainingWei Liu, Kun Qian, Zhenhua Li, Tianyin Xu 等SIGCOMM 2025 · 被引用 8 次
- From Logs to Causal Inference: Diagnosing Large SystemsMarkos Markakis, Brit Youngmann, Trinity Gao, Ziyu Zhang 等VLDB 2025 · 被引用 6 次
- Copper and Wire: Bridging Expressiveness and Performance for Service Mesh PoliciesDivyanshu Saxena, William Zhang, Shankara Pailoor, Isil Dillig 等ASPLOS 2025 · 被引用 3 次
它引用的顶会 Paper4
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo 等ASPLOS 2021 · 被引用 170 次
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 被引用 88 次
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak 等OSDI 2022 · 被引用 38 次
- Scouts: Improving the Diagnosis Process Through Domain-customized Incident RoutingJiaqi Gao, Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri 等SIGCOMM 2020 · 被引用 18 次
相关 Paper
- AutoMAP: Diagnose Your Microservice-based Web Applications AutomaticallyMeng Ma, Jingmin Xu, Yuan Wang, Pengfei Chen 等WWW 2020 · 被引用 144 次
- Look Deep into the Microservice System Anomaly through Very Sparse LogsXinrui Jiang, Yicheng Pan, Meng Ma, Ping WangWWW 2023 · 被引用 16 次
- Closed-loop Network Performance Monitoring and Diagnosis with SpiderMonWeitao Wang, Xinyu Crystal Wu, Praveen Tammana, Ang Chen 等NSDI 2022 · 被引用 37 次
- Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware BenchmarkAoyang Fang, Songhan Zhang, Yifan Yang, Haotong Wu 等FSE 2026 · 被引用 1 次
- Microscope: Queue-based Performance Diagnosis for Network FunctionsJunzhi Gong, Yuliang Li, Bilal Anwer, Aman Shaikh 等SIGCOMM 2020 · 被引用 17 次
