Enhancing Network Failure Mitigation with Performance-Aware Ranking
Pooria Namyar, Arvin Ghavidel, Daniel Crankshaw, Daniel S. Berger, Kevin Hsieh, Srikanth Kandula, Ramesh Govindan, Behnaz Arzani
Abstract
Cloud providers install mitigations to reduce the impact of network failures within their datacenters. Existing network mitigation systems rely on simple local criteria or global proxy metrics to determine the best action. In this paper, we show that we can support a broader range of actions and select more effective mitigations by directly optimizing end-to-end flow-level metrics and analyzing actions holistically. To achieve this, we develop novel techniques to quickly estimate the impact of different mitigations and rank them with high fidelity. Our results on incidents from a large cloud provider show orders of magnitude improvements in flow completion time and throughput. We also show our approach scales to large datacenters.
The author contributed to this work while at Microsoft. 1 We use network-level mitigations and mitigations interchangeably. We briefly discuss application-level mitigations in §3.4.
2 Its auto-mitigation system uses local criteria to determine whether taking a fixed action is better than taking no action.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7c0ece5-74fb-47be-9f26-6885bae8d9f3Cited by top-tier papers4
- Towards LLM-Based Failure Localization in Production-Scale NetworksChenxu Wang, Xumiao Zhang, Runwei Lu, Xianshang Lin et al.SIGCOMM 2025 · 13 citations
- SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud InfrastructuresBo Yang, Huanwu Hu, Yifan Li, Yunguang Li et al.SIGCOMM 2025 · 3 citations
- MoCE: A Mixture-of-Context Aware Experts Framework for Troubleshooting Internet-scale ServicesVipul Harsh, Sayan Sinha, Henry Milner, B. Aditya Prakash et al.NSDI 2026 · 2 citations
- Harp: Improving VPC Network Availability via Efficient Failure Detection and Rerouting in Tencent CloudJiayu Hu, Feng Jin, Xianping Zhou, Kai Zhang et al.NSDI 2026
Builds on15
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh et al.SIGCOMM 2022 · 230 citations
- When Cloud Storage Meets RDMAYixiao Gao, Qiang Li, Lingbo Tang, Yongqing Xi et al.NSDI 2021 · 228 citations
- Expanding across time to deliver bandwidth efficiency and low latencyWilliam M. Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness et al.NSDI 2020 · 194 citations
- OmniMon: Re-architecting Network Telemetry with Resource Efficiency and Full AccuracyQun Huang, Haifeng Sun, Patrick P. C. Lee, Wei Bai et al.SIGCOMM 2020 · 109 citations
- 1RMA: Re-envisioning Remote Memory Access for Multi-tenant DatacentersArjun Singhvi, Aditya Akella, Dan Gibson, Thomas F. Wenisch et al.SIGCOMM 2020 · 70 citations
Related papers
- Optimizing Network Provisioning through CooperationHarsha Sharma, Parth Thakkar, Sagar Bharadwaj, Ranjita Bhagwan et al.NSDI 2022 · 4 citations
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan et al.SIGCOMM 2021 · 63 citations
- Scalable Tail Latency Estimation for Data Center NetworksKevin Zhao, Prateesh Goyal, Mohammad Alizadeh, Thomas E. AndersonNSDI 2023 · 30 citations
- Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM InterruptionsSebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang et al.OSDI 2020 · 35 citations
- ECOTE: Priority-Aware Optical Restoration for WAN Traffic EngineeringYiren Zhao, Kunling He, Zhiquan Wang, Ran Shu et al.EuroSys 2026
