Pilot Execution: Simulating Failure Recovery In Situ for Production Distributed Systems
Zhenyu Li, Angting Cai, Chang Lou
Abstract
Modern distributed systems rely on failure recovery to ensure availability and correctness-ironically, recovery itself often introduces severe and irreversible failures. In this paper, we first study 75 real-world recovery failures to understand common pitfalls in the recovery mechanisms. We find that the challenges primarily arise from cross-component interactions, which are difficult to expose in traditional approaches.
To address this gap, we introduce pilot execution, a new execution model that simulates dry-runs of recovery actions in production distributed systems to enable safe and predictable failure recovery. It enables systems and operators to observe recovery action effects before applying them, reducing the risk of cascading failures and unintended side effects.
We realize pilot execution with PILOT, an analysis framework with a runtime library that makes pilot execution easy to adopt. We evaluate PILOT on five large-scale distributed systems and show that PILOT uncovers 17 out of 20 recovery failures with modest overhead. Our use of PILOT also exposes an unknown recovery bug in the latest version of HBase.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
- Check before You Change: Preventing Correlated Failures in Service UpdatesEnnan Zhai, Ang Chen, Ruzica Piskac, Mahesh Balakrishnan et al.NSDI 2020 · 46 citations
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak et al.OSDI 2022 · 38 citations
- Lessons from the evolution of the Batfish configuration analysis toolMatt Brown, Ari Fogel, Daniel Halperin, Victor Heorhiadi et al.SIGCOMM 2023 · 37 citations
- Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM InterruptionsSebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang et al.OSDI 2020 · 35 citations
Related papers
- Coverage Guided Fault Injection for Cloud SystemsYu Gao, Wensheng Dou, Dong Wang, Wenhan Feng et al.ICSE 2023 · 13 citations
- If At First You Don't Succeed, Try, Try, Again...? Insights and LLM-informed Tooling for Detecting Retry Bugs in Software SystemsBogdan Alexandru Stoica, Utsav Sethi, Yiming Su, Cyrus Zhou et al.SOSP 2024 · 4 citations
- Demystifying and Checking Silent Semantic Violations in Large Distributed SystemsChang Lou, Yuzhuo Jing, Peng HuangOSDI 2022 · 8 citations
- CSnake: Detecting Self-Sustaining Cascading Failure via Causal Stitching of Fault PropagationsShangshu Qian, Lin Tan, Yongle ZhangEuroSys 2026 · 1 citation
- Fractal: Fault-Tolerant Shell-Script DistributionZhicheng Huang, Ramiz Dundar, Yizheng Xie, Konstantinos Kallas et al.NSDI 2026 · 4 citations
