An In-Depth Study of Correlated Failures in Production SSD-Based Data Centers
Shujie Han, Patrick P. C. Lee, Fan Xu, Yi Liu, Cheng He, Jiongzhou Liu
Abstract
Flash-based solid-state drives (SSDs) are increasingly adopted as the mainstream storage media in modern data centers. However, little is known about how SSD failures in the field are correlated, both spatially and temporally. We argue that characterizing correlated failures of SSDs is critical, especially for guiding the design of redundancy protection for high storage reliability. We present an in-depth data-driven analysis on the correlated failures in the SSD-based data centers at Alibaba. We study nearly one million SSDs of 11 drive models based on a dataset of SMART logs, trouble tickets, physical locations, and applications. We show that correlated failures in the same node or rack are common, and study the possible impacting factors on those correlated failures. We also evaluate via trace-driven simulation how various redundancy schemes affect the storage reliability under correlated failures. To this end, we report 15 findings. Our dataset and source code are now released for public use.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Multi-view Feature-based SSD Failure Prediction: What, When, and WhyYuqi Zhang, Wenwen Hao, Ben Niu, Kangkang Liu et al.FAST 2023 · 33 citations
- Perseus: A Fail-Slow Detection Framework for Cloud Storage SystemsRuiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu et al.FAST 2023 · 31 citations
- MSFRD: Mutation Similarity based SSD Failure Rating and Diagnosis for Complex and Volatile Production EnvironmentsYuqi Zhang, Tianyi Zhang, Wenwen Hao, Shuyang Wang et al.USENIX ATC 2024 · 10 citations
- Exploit both SMART Attributes and NAND Flash Wear Characteristics to Effectively Forecast SSD-based Storage Failures in ClustersYunfei Gu, Chentao Wu, Xubin HeUSENIX ATC 2024 · 8 citations
- FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus ScaleWeihao Cui, Ji Zhang, Han Zhao, Chao Liu et al.NSDI 2026 · 7 citations
Builds on2
Related papers
- FailureMiner: A Joint Key Decision Mining Scheme for Practical SSD Failure Prediction and AnalysisShuyang Wang, Yuqi Zhang, Haonan Luo, Kangkang Liu et al.FAST 2026
- Hey Hey, My My, Skewness Is Here to Stay: Challenges and Opportunities in Cloud Block Store TrafficHaonan Wu, Erci Xu, Ligang Wang, Yuandong Hong et al.EuroSys 2025 · 2 citations
- Operational Characteristics of SSDs in Enterprise Storage Systems: A Large-Scale Field StudyStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2022 · 40 citations
- Efficient Bad Block Management with Cluster SimilarityJui-Nan Yen, Yao-Ching Hsieh, Cheng-Yu Chen, Tseng-Yi Chen et al.HPCA 2022 · 10 citations
- POLARDB Meets Computational Storage: Efficiently Support Analytical Workloads in Cloud-Native Relational DatabaseWei Cao, Yang Liu, Zhushi Cheng, Ning Zheng et al.FAST 2020 · 140 citations
