Metastable Failures in the Wild
Lexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak, Rebecca Isaacs, Abutalib Aghayev, Timothy Zhu, Aleksey Charapko
Abstract
Recently, Bronson et al. [7] introduced a framework for understanding a class of failures in distributed systems called metastable failures. The examples of metastable failures presented in that work are simplified versions of failures observed at Facebook. In this work, we study the prevalence of such failures in the wild by scouring over publicly available incident reports from many organizations, ranging from hyperscalers to small companies.
Our main findings are threefold. First, metastable failures are universally observed-we present an in-depth study of 22 metastable failures from 11 different organizations. Second, metastable failures are a recurring pattern in many severe outages-e.g., at least 4 out of 15 major outages in the last decade at Amazon Web Services were caused by metastable failures. Third, we extend the model by Bronson et al. to better reflect the metastable failures seen in the wild by categorizing two types of triggers and two types of amplification mechanisms, which we confirm through developing multiple example applications that reproduce different types of metastable failures in a controlled environment. We believe our work will aid in a deeper understanding of metastable failures and in coming up with solutions to them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 428b0792-7b59-4691-b9d6-a3bbd89a1634Cited by top-tier papers32
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang et al.NSDI 2024 · 192 citations
- On the Limitations of Carbon-Aware Temporal and Spatial Workload Shifting in the CloudThanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin et al.EuroSys 2024 · 73 citations
- On-demand Container Loading in AWS LambdaMarc Brooker, Mike Danilov, Chris Greenwood, Phil PiwonkaUSENIX ATC 2023 · 73 citations
- GL-Cache: Group-level learning for efficient and high-performance cachingJuncheng Yang, Ziming Mao, Yao Yue, K. V. RashmiFAST 2023 · 60 citations
- What's the Story in EBS Glory: Evolutions and Lessons in Building Cloud Block StoreWeidong Zhang, Erci Xu, Qiuping Wang, Xiaolu Zhang et al.FAST 2024 · 37 citations
Builds on3
- Understanding and Detecting Software Upgrade Failures in Distributed SystemsYongle Zhang, Junwen Yang, Zhuqi Jin, Utsav Sethi et al.SOSP 2021 · 40 citations
- Fault-Tolerant Replication with Pull-Based Consensus in MongoDBSiyuan Zhou, Shuai MuNSDI 2021 · 38 citations
- Tolerating Slowdowns in Replicated State Machines using CopilotsKhiem Ngo, Siddhartha Sen, Wyatt LloydOSDI 2020 · 24 citations
Related papers
- Fail through the Cracks: Cross-System Interaction Failures in Modern Cloud SystemsLilia Tang, Chaitanya Bhandari, Yongle Zhang, Anna Karanika et al.EuroSys 2023 · 16 citations
- Vicious Cycles in Distributed Software SystemsShangshu Qian, Wen Fan, Lin Tan, Yongle ZhangASE 2023 · 5 citations
- Cross-System Categorization of Abnormal Traces in Microservice-Based Systems via Meta-LearningYuqing Wang, Mika V. Mäntylä, Serge Demeyer, Mutlu Beyazit et al.FSE 2025
- MicroRes: Versatile Resilience Profiling in Microservices via Degradation Dissemination IndexingTianyi Yang, Cheryl Lee, Jiacheng Shen, Yuxin Su et al.ISSTA 2024 · 5 citations
- Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware BenchmarkAoyang Fang, Songhan Zhang, Yifan Yang, Haotong Wu et al.FSE 2026 · 1 citation
