Zero Downtime Release: Disruption-free Load Balancing of a Multi-Billion User Website
Usama Naseer, Luca Niccolini, Udip Pant, Alan Frindell, Ranjeeth Dasineni, Theophilus A. Benson
Abstract
Modern network infrastructure has evolved into a complex organism to satisfy the performance and availability requirements for the billions of users. Frequent releases such as code upgrades, bug fixes and security updates have become a norm. Millions of globally distributed infrastructure components including servers and load-balancers are restarted frequently from multiple times per-day to per-week. However, every release brings possibilities of disruptions as it can result in reduced cluster capacity, disturb intricate interaction of the components operating at large scales and disrupt the end-users by terminating their connections. The challenge is further complicated by the scale and heterogeneity of supported services and protocols.
In this paper, we leverage different components of the end-toend networking infrastructure to prevent or mask any disruptions in face of releases. Zero Downtime Release is a collection of mechanisms used at Facebook to shield the end-users from any disruptions, preserve the cluster capacity and robustness of the infrastructure when updates are released globally. Our evaluation shows that these mechanisms prevent any significant cluster capacity degradation when a considerable number of productions servers and proxies are restarted and minimizes the disruption for different services (notably TCP, HTTP and publish/subscribe).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa2aae0c-1c81-4bc0-a164-cd266a42bfb3Cited by top-tier papers5
- The ties that un-bind: decoupling IP from web services and sockets for robust addressing agility at CDN-scaleMarwan Fayed, Lorenz Bauer, Vasileios Giotsas, Sami Kerola et al.SIGCOMM 2021 · 23 citations
- Runtime Recovery of Web Applications under Zero-Day ReDoS AttacksZhihao Bai, Ke Wang, Hang Zhu, Yinzhi Cao et al.S&P 2021 · 19 citations
- Rajomon: Decentralized and Coordinated Overload Control for Latency-Sensitive MicroservicesJiali Xing, Akis Giannoukos, Paul Loh, Shuyue Wang et al.NSDI 2025 · 12 citations
- TIPSY: predicting where traffic will ingress a WANMichael Markovitch, Sharad Agarwal, Rodrigo Fonseca, Ryan Beckett et al.SIGCOMM 2022 · 8 citations
- Hermes: Enhancing Layer-7 Cloud Load Balancers with Userspace-Directed I/O Event NotificationTian Pan, Enge Song, Yueshang Zuo, Shaokai Zhang et al.SIGCOMM 2025 · 5 citations
Related papers
- RAS: Continuously Optimized Region-Wide Datacenter Resource AllocationAndrew Newell, Dimitrios Skarlatos, Jingyuan Fan, Pavan Kumar et al.SOSP 2021 · 19 citations
- A Social Network Under Social Distancing: Risk-Driven Backbone Management During COVID-19 and BeyondYiting Xia, Ying Zhang, Zhizhen Zhong, Guanqing Yan et al.NSDI 2021 · 28 citations
- FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production MonitoringDong Young Yoon, Yang Wang, Miao Yu, Elvis Huang et al.SOSP 2024 · 5 citations
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- Capacity-efficient and uncertainty-resilient backbone network planning with hoseSatyajeet Singh Ahuja, Varun Gupta, Vinayak Dangui, Soshant Bali et al.SIGCOMM 2021 · 40 citations
