Understanding, Detecting and Localizing Partial Failures in Large System Software
Chang Lou, Peng Huang, Scott Smith
Abstract
Partial failures occur frequently in cloud systems and can cause serious damage including inconsistency and data loss. Unfortunately, these failures are not well understood. Nor can they be effectively detected. In this paper, we first study 100 real-world partial failures from five mature systems to understand their characteristics. We find that these failures are caused by a variety of defects that require the unique conditions of the production environment to be triggered. Manually writing effective detectors to systematically detect such failures is both time-consuming and error-prone. We thus propose OmegaGen, a static analysis tool that automatically generates customized watchdogs for a given program by using a novel program reduction technique. We have successfully applied OmegaGen to six large distributed systems. In evaluating 22 real-world partial failure cases in these systems, the generated watchdogs can detect 20 cases with a median detection time of 4.2 seconds, and pinpoint the failure scope for 18 cases. The generated watchdogs also expose an unknown, confirmed partial failure bug in the latest version of ZooKeeper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 513ac817-e8ba-46e4-8ce6-09e7d79aece3Cited by top-tier papers28
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang et al.EuroSys 2024 · 175 citations
- Adaptive Performance Anomaly Detection for Online Service Systems via Pattern SketchingZhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang et al.ICSE 2022 · 44 citations
- SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive ValidationYifan Xiong, Yuting Jiang, Ziyue Yang, Lei Qu et al.USENIX ATC 2024 · 41 citations
- Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM InterruptionsSebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang et al.OSDI 2020 · 35 citations
- Perseus: A Fail-Slow Detection Framework for Cloud Storage SystemsRuiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu et al.FAST 2023 · 31 citations
Related papers
- Understanding and Detecting Fail-Slow Hardware Failure Bugs in Cloud SystemsGen Dong, Yu Hua, Yongle Zhang, Zhangyu Chen et al.USENIX ATC 2025 · 6 citations
- Efficient Exposure of Partial Failure Bugs in Distributed Systems with Inferred Abstract StatesHaoze Wu, Jia Pan, Peng HuangNSDI 2024 · 15 citations
- Coverage Guided Fault Injection for Cloud SystemsYu Gao, Wensheng Dou, Dong Wang, Wenhan Feng et al.ICSE 2023 · 13 citations
- Deriving Semantic Checkers from Tests to Detect Silent Failures in Production Distributed SystemsChang Lou, Dimas Shidqi Parikesit, Yujin Huang, Zhewen Yang et al.OSDI 2025 · 6 citations
- Practical Detection of Entropy Loss in Pseudo-Random Number GeneratorsFelix Dörre, Vladimir KlebanovCCS 2016 · 16 citations
