Defcon: Preventing Overload with Graceful Feature Degradation
Justin Meza, Thote Gowda, Ahmed Eid, Tomiwa Ijaware, Dmitry Chernyshev, Yi Yu, Md Nazim Uddin, Rohan Das, Chad Nachiappan, Sari Tran, Shuyang Shi, Tina Luo
Abstract
Every day, billions of people depend on Internet services for communication, commerce, and entertainment. Yet planetaryscale data center infrastructures consisting of millions of servers experience unplanned capacity outages and unexpected demand for resources; how can such infrastructures remain reliable in the face of capacity and workload flux?
In this paper, we introduce Defcon, a system for improving the availability of large-scale, globally-distributed Internet services using graceful feature degradation. In response to overload conditions, Defcon enables site operators to gradually disable less-critical features in order to reduce resource demand. Defcon presents a common interface to product developers to define feature knobs that represent degradation capabilities. Defcon automatically tests knobs to understand each knob's product-and infrastructure-level trade-offs. At Meta, we have used Defcon to improve global product availability in the face of worldwide demand-surges in addition to large-scale infrastructure failures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- TopFull: An Adaptive Top-Down Overload Control for SLO-Oriented MicroservicesJinwoo Park, Jaehyeong Park, Youngmok Jung, Hwijoon Lim et al.SIGCOMM 2024 · 10 citations
- Cooperative Graceful Degradation in Containerized CloudsKapil Agrawal, Sangeetha Abdu JyothiASPLOS 2025 · 2 citations
- Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice InfrastructureMayank Bansal, Milind Chabbi, Kenneth Bogh, Srikanth Prodduturi et al.NSDI 2026 · 1 citation
- CSnake: Detecting Self-Sustaining Cascading Failure via Causal Stitching of Fault PropagationsShangshu Qian, Lin Tan, Yongle ZhangEuroSys 2026 · 1 citation
- ServiceRouter: Hyperscale and Minimal Cost Service Mesh at MetaHarshit Saokar, Soteris Demetriou, Nick Magerko, Max Kontorovich et al.OSDI 2023
Builds on2
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak et al.OSDI 2022 · 38 citations
- Network entitlement: contract-based network sharing with agility and SLO guaranteesSatyajeet Singh Ahuja, Vinayak Dangui, Kirtesh Patil, Manikandan Somasundaram et al.SIGCOMM 2022 · 3 citations
Related papers
- Global Capacity Management With FluxMarius Eriksen, Kaushik Veeraraghavan, Yusuf Abdulghani, Andrew Birchall et al.OSDI 2023 · 9 citations
- Flex: High-Availability Datacenters With Zero Reserved PowerChaojie Zhang, Alok Gautam Kumbhare, Ioannis Manousakis, Deli Zhang et al.ISCA 2021 · 43 citations
- XFaaS: Hyperscale and Low Cost Serverless Functions at MetaAlireza Sahraei, Soteris Demetriou, Amirali Sobhgol, Haoran Zhang et al.SOSP 2023 · 32 citations
- Zero Downtime Release: Disruption-free Load Balancing of a Multi-Billion User WebsiteUsama Naseer, Luca Niccolini, Udip Pant, Alan Frindell et al.SIGCOMM 2020 · 19 citations
- Netcastle: Network Infrastructure Testing At ScaleRob Sherwood, Jinghao Shi, Ying Zhang, Neil Spring et al.NSDI 2024 · 2 citations
