Metastable Failures in the Wild
Lexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak, Rebecca Isaacs, Abutalib Aghayev, Timothy Zhu, Aleksey Charapko
摘要
Recently, Bronson et al. [7] introduced a framework for understanding a class of failures in distributed systems called metastable failures. The examples of metastable failures presented in that work are simplified versions of failures observed at Facebook. In this work, we study the prevalence of such failures in the wild by scouring over publicly available incident reports from many organizations, ranging from hyperscalers to small companies.
Our main findings are threefold. First, metastable failures are universally observed-we present an in-depth study of 22 metastable failures from 11 different organizations. Second, metastable failures are a recurring pattern in many severe outages-e.g., at least 4 out of 15 major outages in the last decade at Amazon Web Services were caused by metastable failures. Third, we extend the model by Bronson et al. to better reflect the metastable failures seen in the wild by categorizing two types of triggers and two types of amplification mechanisms, which we confirm through developing multiple example applications that reproduce different types of metastable failures in a controlled environment. We believe our work will aid in a deeper understanding of metastable failures and in coming up with solutions to them.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang 等NSDI 2024 · 被引用 192 次
- On the Limitations of Carbon-Aware Temporal and Spatial Workload Shifting in the CloudThanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin 等EuroSys 2024 · 被引用 73 次
- On-demand Container Loading in AWS LambdaMarc Brooker, Mike Danilov, Chris Greenwood, Phil PiwonkaUSENIX ATC 2023 · 被引用 73 次
- GL-Cache: Group-level learning for efficient and high-performance cachingJuncheng Yang, Ziming Mao, Yao Yue, K. V. RashmiFAST 2023 · 被引用 60 次
- What's the Story in EBS Glory: Evolutions and Lessons in Building Cloud Block StoreWeidong Zhang, Erci Xu, Qiuping Wang, Xiaolu Zhang 等FAST 2024 · 被引用 37 次
它引用的顶会 Paper3
- Understanding and Detecting Software Upgrade Failures in Distributed SystemsYongle Zhang, Junwen Yang, Zhuqi Jin, Utsav Sethi 等SOSP 2021 · 被引用 40 次
- Fault-Tolerant Replication with Pull-Based Consensus in MongoDBSiyuan Zhou, Shuai MuNSDI 2021 · 被引用 38 次
- Tolerating Slowdowns in Replicated State Machines using CopilotsKhiem Ngo, Siddhartha Sen, Wyatt LloydOSDI 2020 · 被引用 24 次
相关 Paper
- Fail through the Cracks: Cross-System Interaction Failures in Modern Cloud SystemsLilia Tang, Chaitanya Bhandari, Yongle Zhang, Anna Karanika 等EuroSys 2023 · 被引用 16 次
- Vicious Cycles in Distributed Software SystemsShangshu Qian, Wen Fan, Lin Tan, Yongle ZhangASE 2023 · 被引用 5 次
- Cross-System Categorization of Abnormal Traces in Microservice-Based Systems via Meta-LearningYuqing Wang, Mika V. Mäntylä, Serge Demeyer, Mutlu Beyazit 等FSE 2025
- MicroRes: Versatile Resilience Profiling in Microservices via Degradation Dissemination IndexingTianyi Yang, Cheryl Lee, Jiacheng Shen, Yuxin Su 等ISSTA 2024 · 被引用 5 次
- Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware BenchmarkAoyang Fang, Songhan Zhang, Yifan Yang, Haotong Wu 等FSE 2026 · 被引用 1 次
