MoCE: A Mixture-of-Context Aware Experts Framework for Troubleshooting Internet-scale Services
Vipul Harsh, Sayan Sinha, Henry Milner, B. Aditya Prakash, Vyas Sekar, Hui Zhang
摘要
Modern Internet-scale services need to rapidly identify root causes of customer-impacting incidents and remediate them. While there are many algorithms (including LLM-assisted solutions) for root cause analysis, these have significant limitations in terms of coverage, extensibility, and scalability due to the diversity of incidents that can occur at Internet-scale and the complexity of telemetry analysis. We argue the need for a paradigm shift in root cause analysis to depart from algorithm development to a systems approach.
To this end, we introduce a mixture of context-aware experts framework where each "expert" represents a root cause hypothesis exploration. To enable rapid development of new experts and allow computational reuse across the ensemble for scalability, we design an abstraction that allows us to express an expert as a dataflow DAG combining relational, stateful, and statistical operations. To ensure scalability and extensibility, we develop a lazy DAG runtime system that lazily schedules execution of DAG nodes. We implement this idea in MoCE and demonstrate its value using a mix of real-world incident data from four large application analytics providers and synthetically generated incidents. We show that many existing and novel approaches can be expressed succinctly in our framework. We find that MoCE achieves high RCA accuracy (>95%) across diverse incidents compared to 34% for the closest single expert (including prior works) achieving high coverage. We also show the value of the mixture paradigm and the lazy DAG runtime using controlled experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang 等EuroSys 2024 · 被引用 175 次
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo 等ASPLOS 2021 · 被引用 170 次
- Flow Event Telemetry on Programmable Data PlaneYu Zhou, Chen Sun, Hongqiang Harry Liu, Rui Miao 等SIGCOMM 2020 · 被引用 139 次
- FAst in-network GraY failure detection for ISPsEdgar Costa Molero, Stefano Vissicchio, Laurent VanbeverSIGCOMM 2022 · 被引用 26 次
- Scouts: Improving the Diagnosis Process Through Domain-customized Incident RoutingJiaqi Gao, Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri 等SIGCOMM 2020 · 被引用 18 次
相关 Paper
- MetaRCA: A Generalizable Root Cause Analysis Framework for Cloud-Native Systems Powered by Meta Causal KnowledgeShuai Liang, Pengfei Chen, Bozhe Tian, Gou Tan 等FSE 2026 · 被引用 4 次
- Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware BenchmarkAoyang Fang, Songhan Zhang, Yifan Yang, Haotong Wu 等FSE 2026 · 被引用 1 次
- MRCA: Metric-level Root Cause Analysis for Microservices via Multi-Modal DataYidan Wang, Zhouruixing Zhu, Qiuai Fu, Yuchi Ma 等ASE 2024 · 被引用 6 次
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini 等NeurIPS 2022 · 被引用 185 次
- RCAFlow: A Workflow-Informed Hierarchical Planning Multi-Agent System for Root Cause AnalysisYufei Gao, Zhengong Cai, Bowei YangAAAI 2026
