One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed Systems
Ruiming Lu, Yunchi Lu, Yuxuan Jiang, Guangtao Xue, Peng Huang
Abstract
Recent studies have shown that various hardware components exhibit fail-slow behavior at scale. However, the characteristics of distributed software's tolerance of such slow faults remain ill-understood. This paper presents a comprehensive study that investigates the characteristics and current practices of slowfault tolerance in modern distributed software. We focus on the fundamentally nuanced nature of slow faults. We develop a testing pipeline to systematically introduce diverse slow faults, measure their impact under different workloads, and identify the patterns. Our study shows that even small changes can lead to dramatically different reactions. While some systems have added slow-fault handling mechanisms, they are mostly controlled by static thresholds, which can hardly accommodate the highly sensitive and dynamic characteristics. To address this gap, we design ADR, a lightweight library to use within system code and make fail-slow handling adaptive. Evaluation shows ADR significantly reduces the impact of slow faults.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac328d2c-5946-4c8f-89df-403b50178015Cited by top-tier papers4
- Deriving Semantic Checkers from Tests to Detect Silent Failures in Production Distributed SystemsChang Lou, Dimas Shidqi Parikesit, Yujin Huang, Zhewen Yang et al.OSDI 2025 · 6 citations
- kSTEP: Characterization and Deterministic Testing of Linux CPU Scheduler BugsTingjia Cao, Shawn (Wanxiang) Zhong, Caeden Whitaker, Ke Han et al.OSDI 2026 · 1 citation
- Ambulance: Saving BFT through RacingNeil Giridharan, Shubham Mishra, Lorenzo Alvisi, Natacha Crooks et al.OSDI 2026 · 1 citation
- KUBEDIRECT: Unleashing the Full Power of the Cluster Manager for Serverless ComputingSheng Qi, Zhiquan Zhang, Xuanzhe Liu, Xin JinNSDI 2026
Builds on8
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
- Understanding Silent Data Corruptions in a Large Production CPU PopulationShaobu Wang, Guangyan Zhang, Junyu Wei, Yang Wang et al.SOSP 2023 · 56 citations
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak et al.OSDI 2022 · 38 citations
- Perseus: A Fail-Slow Detection Framework for Cloud Storage SystemsRuiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu et al.FAST 2023 · 31 citations
- Toward a Generic Fault Tolerance Technique for Partial Network PartitioningMohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, Samer Al-KiswanyOSDI 2020 · 30 citations
Related papers
- Understanding and Detecting Fail-Slow Hardware Failure Bugs in Cloud SystemsGen Dong, Yu Hua, Yongle Zhang, Zhangyu Chen et al.USENIX ATC 2025 · 6 citations
- A Formal Framework for Predicting Distributed System Performance Under FaultsZiwei Zhou, Si Liu, Zhou Zhou, Peixin Wang et al.FM 2026
- Avicenna: Masking Slowdowns in Replicated State Machines with Counterfactual EvaluationChristopher Hodsdon, Zijian Qin, Khiem Ngo, Siddhartha Sen et al.EuroSys 2026
- Vicious Cycles in Distributed Software SystemsShangshu Qian, Wen Fan, Lin Tan, Yongle ZhangASE 2023 · 5 citations
- CAFault: Enhance Fault Injection Technique in Practical Distributed Systems via Abundant Fault-Dependent ConfigurationsYuanliang Chen, Fuchen Ma, Yuanhang Zhou, Zhen Yan et al.USENIX ATC 2025 · 7 citations
