FailureMiner: A Joint Key Decision Mining Scheme for Practical SSD Failure Prediction and Analysis
Shuyang Wang, Yuqi Zhang, Haonan Luo, Kangkang Liu, Gil Kim, Jongsung Na, Claude Kim, Geunrok Oh, Kyle Choi, Ni Xue, Xing He
摘要
As SSDs become increasingly popular in enterprise data centers, SSD failures have become a key concern for storage system reliability. In this paper, we propose FailureMiner, a joint key decision mining scheme based on SSD monitoring attributes to accurately and clearly identify SSD failure patterns in production environments. First, to address the imbalance between healthy and failed samples caused by the limited number of failed SSDs, FailureMiner introduces selective downsampling to carefully remove non-critical healthy samples, thereby focusing more on the subtle differences between easily confused failure patterns and health patterns. Second, FailureMiner streamlines the decision-making process of the machine learning model in failure prediction by capturing key decision steps based on their joint contribution. By filtering out redundant and noisy information, FailureMiner can capture joint key decisions, i.e., the simplified attribute combinations and value ranges relevant to failures, thus enabling accurate and interpretable identification of failure patterns.
FailureMiner is evaluated on real-world datasets, and the results show that our scheme improves precision and recall by an average of 38.6% and 80.5% respectively, compared with the existing failure prediction methods. The extracted joint key decisions have been deployed in Tencent's data centers to predict failures across more than 350,000 SSDs over a year, enhancing SSD reliability. The joint key decisions also reveal the failure patterns and factors affecting SSD health, which further helps operators handle failures and manufacturers improve product reliability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Making Disk Failure Predictions SMARTer!Sidi Lu, Bing Luo, Tirthak Patel, Yongtao Yao 等FAST 2020 · 被引用 120 次
- Reducing solid-state drive read latency by optimizing read-retryJisung Park, Myungsuk Kim, Myoungjun Chun, Lois Orosa 等ASPLOS 2021 · 被引用 66 次
- Operational Characteristics of SSDs in Enterprise Storage Systems: A Large-Scale Field StudyStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2022 · 被引用 40 次
- Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM InterruptionsSebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang 等OSDI 2020 · 被引用 35 次
- Multi-view Feature-based SSD Failure Prediction: What, When, and WhyYuqi Zhang, Wenwen Hao, Ben Niu, Kangkang Liu 等FAST 2023 · 被引用 33 次
相关 Paper
- An In-Depth Study of Correlated Failures in Production SSD-Based Data CentersShujie Han, Patrick P. C. Lee, Fan Xu, Yi Liu 等FAST 2021 · 被引用 54 次
- MSFRD: Mutation Similarity based SSD Failure Rating and Diagnosis for Complex and Volatile Production EnvironmentsYuqi Zhang, Tianyi Zhang, Wenwen Hao, Shuyang Wang 等USENIX ATC 2024 · 被引用 10 次
- Exploit both SMART Attributes and NAND Flash Wear Characteristics to Effectively Forecast SSD-based Storage Failures in ClustersYunfei Gu, Chentao Wu, Xubin HeUSENIX ATC 2024 · 被引用 8 次
- Perseus: A Fail-Slow Detection Framework for Cloud Storage SystemsRuiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu 等FAST 2023 · 被引用 31 次
- NTAM: Neighborhood-Temporal Attention Model for Disk Failure Prediction in Cloud PlatformsChuan Luo, Pu Zhao, Bo Qiao, Youjiang Wu 等WWW 2021 · 被引用 37 次
