Improving Network Availability with Protective ReRoute
David Wetherall, Abdul Kabbani, Van Jacobson, Jim Winget, Yuchung Cheng, Charles B. Morrey III, Uma Parthavi Moravapalle, Phillipa Gill, Steven Knight, Amin Vahdat
Abstract
We present PRR (Protective ReRoute), a transport technique for shortening user-visible outages that complements routing repair. It can be added to any transport to provide benefits in multipath networks. PRR responds to flow connectivity failure signals, e.g., retransmission timeouts, by changing the FlowLabel on packets of the flow, which causes switches and hosts to choose a different network path that may avoid the outage. To enable it, we shifted our IPv6 network architecture to use the FlowLabel, so that hosts can change the paths of their flows without application involvement. PRR is deployed fleetwide at Google for TCP and Pony Express, where it has been protecting all production traffic for several years. It is also available to our Cloud customers. We find it highly effective for real outages. In a measurement study on our network backbones, adding PRR reduced the cumulative region-pair outage time over TCP with application-level recovery by 63-84%. This is the equivalent of adding 0.4-0.8 "nines" of availability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- μMon: Empowering Microsecond-level Network Monitoring with WaveletsHao Zheng, Chengyuan Huang, Xiangyu Han, Jiaqi Zheng et al.SIGCOMM 2024 · 26 citations
- Unlocking ECMP Programmability for Precise Traffic ControlYadong Liu, Yunming Xiao, Xuan Zhang, Weizhen Dang et al.NSDI 2025 · 10 citations
- CAPA: An Architecture For Operating Cluster Networks With High AvailabilityBingzhe Liu, Colin Scott, Mukarram Tariq, Andrew D. Ferguson et al.NSDI 2024 · 7 citations
- Nezha: SmartNIC-based Virtual Switch Load SharingXing Li, Enge Song, Bowen Yang, Tian Pan et al.SIGCOMM 2025 · 3 citations
- DDoS Detection at the Scale of One Hundred TbpsYunming Xiao, Xijun Luo, Youliang Jiang, Aike Wang et al.NSDI 2026 · 2 citations
Builds on3
- PLB: congestion signals are simple and effective for network load balancingMubashir Adnan Qureshi, Yuchung Cheng, Qianwen Yin, Qiaobin Fu et al.SIGCOMM 2022 · 82 citations
- ARROW: restoration-aware traffic engineeringZhizhen Zhong, Manya Ghobadi, Alaa Khaddaj, Jonathan Leach et al.SIGCOMM 2021 · 37 citations
- Meaningful AvailabilityTamas Hauer, Philipp Hoffmann, John Lunney, Dan Ardelean et al.NSDI 2020 · 20 citations
Related papers
- Staying Alive: Connection Path Reselection at the EdgeRaul Landa, Lorenzo Saino, Lennert Buytenhek, João Taveira AraújoNSDI 2021 · 13 citations
- ZooRoute: Enhancing Cloud-Scale Network Reliability via Candidate Path Provisioning and Overlay Proactive ReroutingXiaoqing Sun, Xing Li, Xionglie Wei, Tian Pan et al.NSDI 2026
- Availability-Aware Routing in Presence of Geographically Correlated FailuresBalázs Vass, Levente Birszki, Erika R. Bérczi-Kovács, Péter Babarczi et al.INFOCOM 2026
- Harp: Improving VPC Network Availability via Efficient Failure Detection and Rerouting in Tencent CloudJiayu Hu, Feng Jin, Xianping Zhou, Kai Zhang et al.NSDI 2026
- PreTE: Traffic Engineering with Predictive FailuresCongcong Miao, Zhizhen Zhong, Yiren Zhao, Arpit Gupta et al.SIGCOMM 2025 · 6 citations
