MoCE: A Mixture-of-Context Aware Experts Framework for Troubleshooting Internet-scale Services
Vipul Harsh, Sayan Sinha, Henry Milner, B. Aditya Prakash, Vyas Sekar, Hui Zhang
Abstract
Modern Internet-scale services need to rapidly identify root causes of customer-impacting incidents and remediate them. While there are many algorithms (including LLM-assisted solutions) for root cause analysis, these have significant limitations in terms of coverage, extensibility, and scalability due to the diversity of incidents that can occur at Internet-scale and the complexity of telemetry analysis. We argue the need for a paradigm shift in root cause analysis to depart from algorithm development to a systems approach.
To this end, we introduce a mixture of context-aware experts framework where each "expert" represents a root cause hypothesis exploration. To enable rapid development of new experts and allow computational reuse across the ensemble for scalability, we design an abstraction that allows us to express an expert as a dataflow DAG combining relational, stateful, and statistical operations. To ensure scalability and extensibility, we develop a lazy DAG runtime system that lazily schedules execution of DAG nodes. We implement this idea in MoCE and demonstrate its value using a mix of real-world incident data from four large application analytics providers and synthetically generated incidents. We show that many existing and novel approaches can be expressed succinctly in our framework. We find that MoCE achieves high RCA accuracy (>95%) across diverse incidents compared to 34% for the closest single expert (including prior works) achieving high coverage. We also show the value of the mixture paradigm and the lazy DAG runtime using controlled experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang et al.EuroSys 2024 · 175 citations
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo et al.ASPLOS 2021 · 170 citations
- Flow Event Telemetry on Programmable Data PlaneYu Zhou, Chen Sun, Hongqiang Harry Liu, Rui Miao et al.SIGCOMM 2020 · 139 citations
- FAst in-network GraY failure detection for ISPsEdgar Costa Molero, Stefano Vissicchio, Laurent VanbeverSIGCOMM 2022 · 26 citations
- Scouts: Improving the Diagnosis Process Through Domain-customized Incident RoutingJiaqi Gao, Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri et al.SIGCOMM 2020 · 18 citations
Related papers
- MetaRCA: A Generalizable Root Cause Analysis Framework for Cloud-Native Systems Powered by Meta Causal KnowledgeShuai Liang, Pengfei Chen, Bozhe Tian, Gou Tan et al.FSE 2026 · 4 citations
- Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware BenchmarkAoyang Fang, Songhan Zhang, Yifan Yang, Haotong Wu et al.FSE 2026 · 1 citation
- MRCA: Metric-level Root Cause Analysis for Microservices via Multi-Modal DataYidan Wang, Zhouruixing Zhu, Qiuai Fu, Yuchi Ma et al.ASE 2024 · 6 citations
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini et al.NeurIPS 2022 · 185 citations
- RCAFlow: A Workflow-Informed Hierarchical Planning Multi-Agent System for Root Cause AnalysisYufei Gao, Zhengong Cai, Bowei YangAAAI 2026
