Scouts: Improving the Diagnosis Process Through Domain-customized Incident Routing
Jiaqi Gao, Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri, Vincent Liu, Ricardo Bianchini, Ramaswamy Aditya, Xiaohang Wang, Henry Lee, David A. Maltz, Minlan Yu, Behnaz Arzani
Abstract
Incident routing is critical for maintaining service level objectives in the cloud: the time-to-diagnosis can increase by 10x due to mis-routings. Properly routing incidents is challenging because of the complexity of today's data center (DC) applications and their dependencies. For instance, an application running on a VM might rely on a functioning host-server, remote-storage service, and virtual and physical network components. It is hard for any one team, rule-based system, or even machine learning solution to fully learn the complexity and solve the incident routing problem. We propose a different approach using per-team Scouts. Each teams' Scout acts as its gate-keeper --- it routes relevant incidents to the team and routes-away unrelated ones. We solve the problem through a collection of these Scouts. Our PhyNet Scout alone --- currently deployed in production --- reduces the time-to-mitigation of 65% of mis-routed incidents in our dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Minder: Faulty Machine Detection for Large-scale Distributed Model TrainingYangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang et al.NSDI 2025 · 36 citations
- Graph-based Incident Aggregation for Large-Scale Online Service SystemsZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang et al.ASE 2021 · 29 citations
- NetAssistant: Dialogue Based Network Diagnosis in Data Center NetworksHaopei Wang, Anubhavnidhi Abhashkumar, Changyu Lin, Tianrong Zhang et al.NSDI 2024 · 21 citations
- Murphy: Performance Diagnosis of Distributed Cloud ApplicationsVipul Harsh, Wenxuan Zhou, Sachin Ashok, Radhika Niranjan Mysore et al.SIGCOMM 2023 · 16 citations
- Towards LLM-Based Failure Localization in Production-Scale NetworksChenxu Wang, Xumiao Zhang, Runwei Lu, Xianshang Lin et al.SIGCOMM 2025 · 13 citations
Builds on1
Related papers
- Triangle: Empowering Incident Triage with Multi-AgentZhaoyang Yu, Aoyang Fang, Minghua Ma, Jaskaran Singh Walia et al.ASE 2025 · 2 citations
- Scalable Tail Latency Estimation for Data Center NetworksKevin Zhao, Prateesh Goyal, Mohammad Alizadeh, Thomas E. AndersonNSDI 2023 · 30 citations
- SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud InfrastructuresBo Yang, Huanwu Hu, Yifan Li, Yunguang Li et al.SIGCOMM 2025 · 3 citations
- Fighting the Fog of War: Automated Incident Detection for Cloud SystemsLiqun Li, Xu Zhang, Xin Zhao, Hongyu Zhang et al.USENIX ATC 2021 · 43 citations
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan et al.SIGCOMM 2021 · 63 citations
