Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure
Mayank Bansal, Milind Chabbi, Kenneth Bogh, Srikanth Prodduturi, Kevin Xu, Amit Kumar, David Bell, Ranjib Dey, Yufei Ren, Sachin Sharma, Juan Marcano, Shriniket Kale
Abstract
Operating a global, real-time platform at Uber's scale requires infrastructure that is both resilient and cost-efficient. Historically, reliability was ensured through a costly 2x capacity model--each service provisioned to handle global traffic independently across two regions--leaving half the fleet idle. We present Uber's Failover Architecture (UFA), which replaces the uniform 2x model with a differentiated architecture aligned to business criticality. Critical services retain failover guarantees, while non-critical services opportunistically use failover buffer capacity reserved for critical services during steady state. During rare"full-peak"failovers, non-critical services are selectively preempted and rapidly restored, with differentiated Service-Level Agreements (SLAs) using on-demand capacity. Automated safeguards, including dependency analysis and regression gates, ensure critical services continue to function even while non-critical services are unavailable. The quantitative impact is significant: UFA reduces steady-state provisioning from 2x to 1.3x, raising utilization from 20% to 30% while sustaining 99.97% availability. To date, UFA has hardened over 4,000 unsafe dependencies, eliminated over one million CPU cores from a baseline of about four million cores.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b1ec6cf-3127-4de2-abed-d0ef8f84213dBuilds on10
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
- Check before You Change: Preventing Correlated Failures in Service UpdatesEnnan Zhai, Ang Chen, Ruzica Piskac, Mahesh Balakrishnan et al.NSDI 2020 · 46 citations
- Ripple: Profile-Guided Instruction Cache Replacement for Data Center ApplicationsTanvir Ahmed Khan, Dexin Zhang, Akshitha Sriraman, Joseph Devietti et al.ISCA 2021 · 33 citations
- Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsTanvir Ahmed Khan, Nathan Brown, Akshitha Sriraman, Niranjan K. Soundararajan et al.MICRO 2021 · 33 citations
- Thermometer: profile-guided btb replacement for data center applicationsShixin Song, Tanvir Ahmed Khan, Sara Mahdizadeh-Shahri, Akshitha Sriraman et al.ISCA 2022 · 23 citations
Related papers
- XFaaS: Hyperscale and Low Cost Serverless Functions at MetaAlireza Sahraei, Soteris Demetriou, Amirali Sobhgol, Haoran Zhang et al.SOSP 2023 · 32 citations
- RPCS hield : Defending Microservices Against Cascading FailuresMilind Chabbi, Sonal Mahajan, Ivan Beschastnikh, René Just et al.SOSP 2026
- Grad: Intelligent Microservice Scaling by Harnessing Resource FungibilityLiao Chen, Chenyu Lin, Shutian Luo, Huanle Xu et al.HPCA 2025 · 5 citations
- Flex: High-Availability Datacenters With Zero Reserved PowerChaojie Zhang, Alok Gautam Kumbhare, Ioannis Manousakis, Deli Zhang et al.ISCA 2021 · 43 citations
- Live in the Express LanePatrick Jahnke, Vincent Riesop, Pierre-Louis Roman, Pavel Chuprikov et al.USENIX ATC 2021 · 7 citations
