Understanding and Detecting Software Upgrade Failures in Distributed Systems
Yongle Zhang, Junwen Yang, Zhuqi Jin, Utsav Sethi, Kirk Rodrigues, Shan Lu, Ding Yuan
Abstract
Upgrade is one of the most disruptive yet unavoidable maintenance tasks that undermine the availability of distributed systems. Any failure during an upgrade is catastrophic, as it further extends the service disruption caused by the upgrade. The increasing adoption of continuous deployment further increases the frequency and burden of the upgrade task. In practice, upgrade failures have caused many of today's highprofile cloud outages. Unfortunately, there has been little understanding of their characteristics.
This paper presents an in-depth study of 123 real-world upgrade failures that were previously reported by users in 8 widely used distributed systems, shedding lights on the severity, root causes, exposing conditions, and fix strategies of upgrade failures. Guided by our study, we have designed a testing framework DUPTester that revealed 20 previously unknown upgrade failures in 4 distributed systems, and applied a series of static checkers DUPChecker that discovered over 800 cross-version data-format incompatibilities that can lead to upgrade failures. DUPChecker has been requested by HBase developers to be integrated into their toolchain.
• Computer systems organization → Availability; • Software and its engineering → Software testing and debugging; Automated static analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 596fff35-88a3-415f-b651-3c1f2949bc8cCited by top-tier papers23
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang et al.EuroSys 2024 · 175 citations
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann et al.ICSE 2023 · 93 citations
- Automatic Reliability Testing For Cluster Management ControllersXudong Sun, Wenqing Luo, Jiawei Tyler Gu, Aishwarya Ganesan et al.OSDI 2022 · 44 citations
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak et al.OSDI 2022 · 38 citations
- Model Checking Guided Testing for Distributed SystemsDong Wang, Wensheng Dou, Yu Gao, Chenao Wu et al.EuroSys 2023 · 21 citations
Builds on1
Related papers
- UpFuzz: Detecting Data Format Incompatibility Bugs during Distributed Storage System UpgradeKe Han, P. C. Sruthi, Yayu Wang, Yaoxu Song et al.NSDI 2026 · 3 citations
- Demystifying and Checking Silent Semantic Violations in Large Distributed SystemsChang Lou, Yuzhuo Jing, Peng HuangOSDI 2022 · 8 citations
- Efficient Exposure of Partial Failure Bugs in Distributed Systems with Inferred Abstract StatesHaoze Wu, Jia Pan, Peng HuangNSDI 2024 · 15 citations
- Deriving Semantic Checkers from Tests to Detect Silent Failures in Production Distributed SystemsChang Lou, Dimas Shidqi Parikesit, Yujin Huang, Zhewen Yang et al.OSDI 2025 · 6 citations
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
