Vicious Cycles in Distributed Software Systems
Shangshu Qian, Wen Fan, Lin Tan, Yongle Zhang
摘要
A major threat to distributed software systems' reliability is vicious cycles, which are observed when an event in the distributed software system's execution causes a system degradation, and the degradation, in turn, causes more of such events. Vicious cycles often result in large-scale cloud outages that are hard to recover from due to their self-reinforcing nature. This paper formally defines Vicious Cycle, and conducts the first in-depth study of 33 real-world vicious cycles in 13 widely-used open-source distributed software systems, shedding light on the root causes, triggering conditions, and fixing strategies of vicious cycles, with over a dozen concrete implications to combat them. Our findings show that the majority of the vicious cycles are caused by incorrect error handlers, where the handlers do not obtain enough information to distinguish between 1) an error induced by incoming requests and 2) an error induced by an unexpected interference from another error handler. This paper further performs a feasibility study by 1) building a monitoring tool that prevents one type of vicious cycle by collecting information to make a more informed decision in error handling, and 2) investigating the effectiveness of one commonly suggested practice-injecting exponential backoff-to prevent vicious cycles induced by unconstrained retry.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Deriving Semantic Checkers from Tests to Detect Silent Failures in Production Distributed SystemsChang Lou, Dimas Shidqi Parikesit, Yujin Huang, Zhewen Yang 等OSDI 2025 · 被引用 6 次
- CSnake: Detecting Self-Sustaining Cascading Failure via Causal Stitching of Fault PropagationsShangshu Qian, Lin Tan, Yongle ZhangEuroSys 2026 · 被引用 1 次
- Pilot Execution: Simulating Failure Recovery In Situ for Production Distributed SystemsZhenyu Li, Angting Cai, Chang LouNSDI 2026
它引用的顶会 Paper6
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 被引用 88 次
- Understanding and Detecting Software Upgrade Failures in Distributed SystemsYongle Zhang, Junwen Yang, Zhuqi Jin, Utsav Sethi 等SOSP 2021 · 被引用 40 次
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak 等OSDI 2022 · 被引用 38 次
- Toward a Generic Fault Tolerance Technique for Partial Network PartitioningMohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, Samer Al-KiswanyOSDI 2020 · 被引用 30 次
- RESIN: A Holistic Service for Dealing with Memory Leaks in Production Cloud InfrastructureChang Lou, Cong Chen, Peng Huang, Yingnong Dang 等OSDI 2022 · 被引用 18 次
相关 Paper
- Uncovering Similar but Different Packages in PyPI and Potential Security ThreatsSunha Park, Soojin Han, Seunghoon WooFSE 2026
- One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed SystemsRuiming Lu, Yunchi Lu, Yuxuan Jiang, Guangtao Xue 等NSDI 2025 · 被引用 14 次
- If At First You Don't Succeed, Try, Try, Again...? Insights and LLM-informed Tooling for Detecting Retry Bugs in Software SystemsBogdan Alexandru Stoica, Utsav Sethi, Yiming Su, Cyrus Zhou 等SOSP 2024 · 被引用 4 次
- Fixing dependency errors for Python build reproducibilitySuchita Mukherjee, Abigail Almanza, Cindy Rubio-GonzálezISSTA 2021 · 被引用 55 次
- Understanding and Detecting Peer Dependency Resolving Loop in npm EcosystemXingyu Wang, Mingsen Wang, Wenbo Shen, Rui ChangICSE 2025 · 被引用 1 次
