SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud Infrastructures
Bo Yang, Huanwu Hu, Yifan Li, Yunguang Li, Xiangyu Tang, Bingchuan Tian, Gongwei Wu, Jianfeng Xu, Xumiao Zhang, Feng Chen, Cheng Wang, Ennan Zhai
Abstract
For providers operating large-scale global networks, the timeliness of network failure recovery significantly affects the reliability of network services. Ideally, a network monitoring system should have enough coverage to detect even minor issues, but high coverage means alert floods during severe network failures. In practice, there is a gap between the flooding raw alerts data collected by network monitoring tools and the readable information needed for failure diagnosis. Existing solutions using limited network monitoring data sources and heuristic diagnostic rules, lack comprehensive coverage and the capability to address severe failures, especially which network operators have never handled a similar one before. This paper presents SkyNet, a network analysis system to extract scope and severity information from alert floods. SkyNet ensures comprehensive coverage by integrating multiple monitoring data sources through a uniform input format, enhancing extensibility for new network monitoring tools. During alert floods, SkyNet groups alerts, assesses their severity, and filters out insignificant ones to aid network operators in mitigating network failures. To date, SkyNet has been running stably on our network for one and a half years without any false negatives and has successfully reduced the time-to-mitigation for over 80% of network failures since its deployment in production.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb5f9fc2-9534-445b-936f-9b48298a1e7bCited by top-tier papers2
- CrossCheck: Input Validation for WAN Control SystemsAlexander Krentsel, Rishabh Iyer, Isaac Keslassy, Bharath Modhipalli et al.NSDI 2026 · 2 citations
- Harp: Improving VPC Network Availability via Efficient Failure Detection and Rerouting in Tencent CloudJiayu Hu, Feng Jin, Xianping Zhou, Kai Zhang et al.NSDI 2026
Builds on16
- LongRoPE: Extending LLM Context Window Beyond 2 Million TokensYiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu et al.ICML 2024 · 316 citations
- PINT: Probabilistic In-band Network TelemetryRan Ben Basat, Sivaramakrishnan Ramanathan, Yuliang Li, Gianni Antichi et al.SIGCOMM 2020 · 268 citations
- NetLLM: Adapting Large Language Models for NetworkingDuo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang et al.SIGCOMM 2024 · 162 citations
- OmniMon: Re-architecting Network Telemetry with Resource Efficiency and Full AccuracyQun Huang, Haifeng Sun, Patrick P. C. Lee, Wei Bai et al.SIGCOMM 2020 · 109 citations
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann et al.ICSE 2023 · 93 citations
Related papers
- Skyline: A Cloud Centric Internet Monitoring EngineShixian Guo, Ziqian Liu, Yangyang Bai, Yuan Chen et al.NSDI 2026 · 4 citations
- Scouts: Improving the Diagnosis Process Through Domain-customized Incident RoutingJiaqi Gao, Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri et al.SIGCOMM 2020 · 18 citations
- Tomography-based progressive network recovery and critical service restoration after massive failuresViviana Arrigoni, Matteo Prata, Novella BartoliniINFOCOM 2023 · 3 citations
- Flow Event Telemetry on Programmable Data PlaneYu Zhou, Chen Sun, Hongqiang Harry Liu, Rui Miao et al.SIGCOMM 2020 · 139 citations
- Enhancing Network Failure Mitigation with Performance-Aware RankingPooria Namyar, Arvin Ghavidel, Daniel Crankshaw, Daniel S. Berger et al.NSDI 2025
