Scouts: Improving the Diagnosis Process Through Domain-customized Incident Routing
Jiaqi Gao, Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri, Vincent Liu, Ricardo Bianchini, Ramaswamy Aditya, Xiaohang Wang, Henry Lee, David A. Maltz, Minlan Yu, Behnaz Arzani
摘要
Incident routing is critical for maintaining service level objectives in the cloud: the time-to-diagnosis can increase by 10x due to mis-routings. Properly routing incidents is challenging because of the complexity of today's data center (DC) applications and their dependencies. For instance, an application running on a VM might rely on a functioning host-server, remote-storage service, and virtual and physical network components. It is hard for any one team, rule-based system, or even machine learning solution to fully learn the complexity and solve the incident routing problem. We propose a different approach using per-team Scouts. Each teams' Scout acts as its gate-keeper --- it routes relevant incidents to the team and routes-away unrelated ones. We solve the problem through a collection of these Scouts. Our PhyNet Scout alone --- currently deployed in production --- reduces the time-to-mitigation of 65% of mis-routed incidents in our dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Minder: Faulty Machine Detection for Large-scale Distributed Model TrainingYangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang 等NSDI 2025 · 被引用 36 次
- Graph-based Incident Aggregation for Large-Scale Online Service SystemsZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang 等ASE 2021 · 被引用 29 次
- NetAssistant: Dialogue Based Network Diagnosis in Data Center NetworksHaopei Wang, Anubhavnidhi Abhashkumar, Changyu Lin, Tianrong Zhang 等NSDI 2024 · 被引用 21 次
- Murphy: Performance Diagnosis of Distributed Cloud ApplicationsVipul Harsh, Wenxuan Zhou, Sachin Ashok, Radhika Niranjan Mysore 等SIGCOMM 2023 · 被引用 16 次
- Towards LLM-Based Failure Localization in Production-Scale NetworksChenxu Wang, Xumiao Zhang, Runwei Lu, Xianshang Lin 等SIGCOMM 2025 · 被引用 13 次
它引用的顶会 Paper1
相关 Paper
- Triangle: Empowering Incident Triage with Multi-AgentZhaoyang Yu, Aoyang Fang, Minghua Ma, Jaskaran Singh Walia 等ASE 2025 · 被引用 2 次
- Scalable Tail Latency Estimation for Data Center NetworksKevin Zhao, Prateesh Goyal, Mohammad Alizadeh, Thomas E. AndersonNSDI 2023 · 被引用 30 次
- SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud InfrastructuresBo Yang, Huanwu Hu, Yifan Li, Yunguang Li 等SIGCOMM 2025 · 被引用 3 次
- Fighting the Fog of War: Automated Incident Detection for Cloud SystemsLiqun Li, Xu Zhang, Xin Zhao, Hongyu Zhang 等USENIX ATC 2021 · 被引用 43 次
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan 等SIGCOMM 2021 · 被引用 63 次
