Perseus: A Fail-Slow Detection Framework for Cloud Storage Systems
Ruiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu, Zhaosheng Zhu, Mengtian Wang, Zongpeng Zhu, Guangtao Xue, Jiwu Shu, Minglu Li, Jiesheng Wu
Abstract
The newly-emerging "fail-slow" failures plague both software and hardware where the victim components are still functioning yet with degraded performance. To address this problem, this paper presents PERSEUS, a practical fail-slow detection framework for storage devices. PERSEUS leverages a light regression-based model to fast pinpoint and analyze fail-slow failures at the granularity of drives. Within a 10month close monitoring on 248K drives, PERSEUS managed to find 304 fail-slow cases. Isolating them can reduce the (node-level) 99.99 th tail latency by 48%. We assemble a large-scale fail-slow dataset (including 41K normal drives and 315 verified fail-slow drives) from our production traces, based on which we provide root cause analysis on fail-slow drives covering a variety of ill-implemented scheduling, hardware defects, and environmental factors. We have released the dataset to the public for fail-slow study.
than 300 fail-slow drives. By isolating and/or replacing the identified fail-slow drives, we significantly reduce the nodelevel tail latency. The 95 th , 99 th , and 99.99 th write latencies drop by 31%, 46%, and 48%, respectively.
We compare PERSEUS to previous fail-slow detection methods as follows. We assemble a large-scale fail-slow dataset (including 315 verified fail-slow drives and around 41K of their cluster-wise peer drives) from our production traces, and build a test benchmark based on the dataset. The benchmark evaluations indicate that PERSEUS outperforms all previous methods, achieving a precision of 0.99 and a recall of 1.00. We also evaluate the effectiveness of components and the sensitivity of parameters in PERSEUS. The results show that PERSEUS can serve as a non-intrusive (based on monitoring traces), fine-grained (per-drive), general (one set of parameters fits all) and accurate (high precision and recall) fail-slow detection framework for the cloud storage systems.
We have also analyzed the reasons for fail-slow failures and discover a wide variety of root causes including illimplemented scheduling (e.g., unnecessary resource contention), hardware flaws (e.g., bad sectors for HDDs), and environmental factors (e.g., temperature and power).
This paper makes the following contributions.
• We share our lessons on detecting fail-slow failures in largescale data centers from three unsuccessful attempts.
• We propose the design of PERSEUS, a non-intrusive, finegrained and general fail-slow detection framework.
• We assemble a large-scale fail-slow dataset 1 and build a fail-slow test benchmark.
• We provide an in-depth root cause analysis of fail-slow failures from the perspective of various factors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76acdf6a-0e84-42f6-82cc-542154cf447dCited by top-tier papers16
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang et al.USENIX ATC 2025 · 19 citations
- Burstable Cloud Block Storage with Data Processing UnitsJunyi Shu, Kun Qian, Ennan Zhai, Xuanzhe Liu et al.OSDI 2024 · 17 citations
- Efficient Exposure of Partial Failure Bugs in Distributed Systems with Inferred Abstract StatesHaoze Wu, Jia Pan, Peng HuangNSDI 2024 · 15 citations
- One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed SystemsRuiming Lu, Yunchi Lu, Yuxuan Jiang, Guangtao Xue et al.NSDI 2025 · 14 citations
- Chronos: Finding Timeout Bugs in Practical Distributed Systems by Deep-Priority Fuzzing with Transient DelayYuanliang Chen, Fuchen Ma, Yuanhang Zhou, Ming Gu et al.S&P 2024 · 12 citations
Builds on4
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
- A Study of SSD Reliability in Large Scale Enterprise Storage DeploymentsStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2020 · 69 citations
- An In-Depth Study of Correlated Failures in Production SSD-Based Data CentersShujie Han, Patrick P. C. Lee, Fan Xu, Yi Liu et al.FAST 2021 · 54 citations
- Understanding and dealing with hard faults in persistent memory systemsBrian Choi, Randal C. Burns, Peng HuangEuroSys 2021 · 11 citations
Related papers
- Understanding and Detecting Fail-Slow Hardware Failure Bugs in Cloud SystemsGen Dong, Yu Hua, Yongle Zhang, Zhangyu Chen et al.USENIX ATC 2025 · 6 citations
- Come Hell or Still Water: Alleviating Tail Latency in Cloud Block StoreChaolei Hu, Kun Qian, Erci Xu, Yifan Shen et al.NSDI 2026
- Scheduling Cloud Block Storage Proactively and Reactively with OmarXinqi Chen, Weidong Zhang, Zhongyu Wang, Erci Xu et al.EuroSys 2026
- ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production TestingMike Chow, Yang Wang, William Wang, Ayichew Hailu et al.OSDI 2024 · 11 citations
- Making Disk Failure Predictions SMARTer!Sidi Lu, Bing Luo, Tirthak Patel, Yongtao Yao et al.FAST 2020 · 120 citations
