One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed Systems
Ruiming Lu, Yunchi Lu, Yuxuan Jiang, Guangtao Xue, Peng Huang
摘要
Recent studies have shown that various hardware components exhibit fail-slow behavior at scale. However, the characteristics of distributed software's tolerance of such slow faults remain ill-understood. This paper presents a comprehensive study that investigates the characteristics and current practices of slowfault tolerance in modern distributed software. We focus on the fundamentally nuanced nature of slow faults. We develop a testing pipeline to systematically introduce diverse slow faults, measure their impact under different workloads, and identify the patterns. Our study shows that even small changes can lead to dramatically different reactions. While some systems have added slow-fault handling mechanisms, they are mostly controlled by static thresholds, which can hardly accommodate the highly sensitive and dynamic characteristics. To address this gap, we design ADR, a lightweight library to use within system code and make fail-slow handling adaptive. Evaluation shows ADR significantly reduces the impact of slow faults.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Deriving Semantic Checkers from Tests to Detect Silent Failures in Production Distributed SystemsChang Lou, Dimas Shidqi Parikesit, Yujin Huang, Zhewen Yang 等OSDI 2025 · 被引用 6 次
- kSTEP: Characterization and Deterministic Testing of Linux CPU Scheduler BugsTingjia Cao, Shawn (Wanxiang) Zhong, Caeden Whitaker, Ke Han 等OSDI 2026 · 被引用 1 次
- Ambulance: Saving BFT through RacingNeil Giridharan, Shubham Mishra, Lorenzo Alvisi, Natacha Crooks 等OSDI 2026 · 被引用 1 次
- KUBEDIRECT: Unleashing the Full Power of the Cluster Manager for Serverless ComputingSheng Qi, Zhiquan Zhang, Xuanzhe Liu, Xin JinNSDI 2026
它引用的顶会 Paper8
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 被引用 88 次
- Understanding Silent Data Corruptions in a Large Production CPU PopulationShaobu Wang, Guangyan Zhang, Junyu Wei, Yang Wang 等SOSP 2023 · 被引用 56 次
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak 等OSDI 2022 · 被引用 38 次
- Perseus: A Fail-Slow Detection Framework for Cloud Storage SystemsRuiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu 等FAST 2023 · 被引用 31 次
- Toward a Generic Fault Tolerance Technique for Partial Network PartitioningMohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, Samer Al-KiswanyOSDI 2020 · 被引用 30 次
相关 Paper
- Understanding and Detecting Fail-Slow Hardware Failure Bugs in Cloud SystemsGen Dong, Yu Hua, Yongle Zhang, Zhangyu Chen 等USENIX ATC 2025 · 被引用 6 次
- A Formal Framework for Predicting Distributed System Performance Under FaultsZiwei Zhou, Si Liu, Zhou Zhou, Peixin Wang 等FM 2026
- Avicenna: Masking Slowdowns in Replicated State Machines with Counterfactual EvaluationChristopher Hodsdon, Zijian Qin, Khiem Ngo, Siddhartha Sen 等EuroSys 2026
- Vicious Cycles in Distributed Software SystemsShangshu Qian, Wen Fan, Lin Tan, Yongle ZhangASE 2023 · 被引用 5 次
- CAFault: Enhance Fault Injection Technique in Practical Distributed Systems via Abundant Fault-Dependent ConfigurationsYuanliang Chen, Fuchen Ma, Yuanhang Zhou, Zhen Yan 等USENIX ATC 2025 · 被引用 7 次
