Perseus: A Fail-Slow Detection Framework for Cloud Storage Systems
Ruiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu, Zhaosheng Zhu, Mengtian Wang, Zongpeng Zhu, Guangtao Xue, Jiwu Shu, Minglu Li, Jiesheng Wu
摘要
The newly-emerging "fail-slow" failures plague both software and hardware where the victim components are still functioning yet with degraded performance. To address this problem, this paper presents PERSEUS, a practical fail-slow detection framework for storage devices. PERSEUS leverages a light regression-based model to fast pinpoint and analyze fail-slow failures at the granularity of drives. Within a 10month close monitoring on 248K drives, PERSEUS managed to find 304 fail-slow cases. Isolating them can reduce the (node-level) 99.99 th tail latency by 48%. We assemble a large-scale fail-slow dataset (including 41K normal drives and 315 verified fail-slow drives) from our production traces, based on which we provide root cause analysis on fail-slow drives covering a variety of ill-implemented scheduling, hardware defects, and environmental factors. We have released the dataset to the public for fail-slow study.
than 300 fail-slow drives. By isolating and/or replacing the identified fail-slow drives, we significantly reduce the nodelevel tail latency. The 95 th , 99 th , and 99.99 th write latencies drop by 31%, 46%, and 48%, respectively.
We compare PERSEUS to previous fail-slow detection methods as follows. We assemble a large-scale fail-slow dataset (including 315 verified fail-slow drives and around 41K of their cluster-wise peer drives) from our production traces, and build a test benchmark based on the dataset. The benchmark evaluations indicate that PERSEUS outperforms all previous methods, achieving a precision of 0.99 and a recall of 1.00. We also evaluate the effectiveness of components and the sensitivity of parameters in PERSEUS. The results show that PERSEUS can serve as a non-intrusive (based on monitoring traces), fine-grained (per-drive), general (one set of parameters fits all) and accurate (high precision and recall) fail-slow detection framework for the cloud storage systems.
We have also analyzed the reasons for fail-slow failures and discover a wide variety of root causes including illimplemented scheduling (e.g., unnecessary resource contention), hardware flaws (e.g., bad sectors for HDDs), and environmental factors (e.g., temperature and power).
This paper makes the following contributions.
• We share our lessons on detecting fail-slow failures in largescale data centers from three unsuccessful attempts.
• We propose the design of PERSEUS, a non-intrusive, finegrained and general fail-slow detection framework.
• We assemble a large-scale fail-slow dataset 1 and build a fail-slow test benchmark.
• We provide an in-depth root cause analysis of fail-slow failures from the perspective of various factors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang 等USENIX ATC 2025 · 被引用 19 次
- Burstable Cloud Block Storage with Data Processing UnitsJunyi Shu, Kun Qian, Ennan Zhai, Xuanzhe Liu 等OSDI 2024 · 被引用 17 次
- Efficient Exposure of Partial Failure Bugs in Distributed Systems with Inferred Abstract StatesHaoze Wu, Jia Pan, Peng HuangNSDI 2024 · 被引用 15 次
- One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed SystemsRuiming Lu, Yunchi Lu, Yuxuan Jiang, Guangtao Xue 等NSDI 2025 · 被引用 14 次
- Chronos: Finding Timeout Bugs in Practical Distributed Systems by Deep-Priority Fuzzing with Transient DelayYuanliang Chen, Fuchen Ma, Yuanhang Zhou, Ming Gu 等S&P 2024 · 被引用 12 次
它引用的顶会 Paper4
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 被引用 88 次
- A Study of SSD Reliability in Large Scale Enterprise Storage DeploymentsStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2020 · 被引用 69 次
- An In-Depth Study of Correlated Failures in Production SSD-Based Data CentersShujie Han, Patrick P. C. Lee, Fan Xu, Yi Liu 等FAST 2021 · 被引用 54 次
- Understanding and dealing with hard faults in persistent memory systemsBrian Choi, Randal C. Burns, Peng HuangEuroSys 2021 · 被引用 11 次
相关 Paper
- Understanding and Detecting Fail-Slow Hardware Failure Bugs in Cloud SystemsGen Dong, Yu Hua, Yongle Zhang, Zhangyu Chen 等USENIX ATC 2025 · 被引用 6 次
- Come Hell or Still Water: Alleviating Tail Latency in Cloud Block StoreChaolei Hu, Kun Qian, Erci Xu, Yifan Shen 等NSDI 2026
- Scheduling Cloud Block Storage Proactively and Reactively with OmarXinqi Chen, Weidong Zhang, Zhongyu Wang, Erci Xu 等EuroSys 2026
- ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production TestingMike Chow, Yang Wang, William Wang, Ayichew Hailu 等OSDI 2024 · 被引用 11 次
- Making Disk Failure Predictions SMARTer!Sidi Lu, Bing Luo, Tirthak Patel, Yongtao Yao 等FAST 2020 · 被引用 120 次
