USENIX ATC2020顶会
HDDse: Enabling High-Dimensional Disk State Embedding for Generic Failure Detection System of Heterogeneous Disks in Large Data Centers
Ji Zhang, Ping Huang, Ke Zhou, Ming Xie, Sebastian Schelter
摘要
The reliability of a storage system is crucial in large data centers. Hard disks are widely used as primary storage devices in modern data centers, where disk failures constantly happen. Disk failures could lead to a serious system interrupt or even permanent data loss. Many hard disk failure detection approaches have been proposed to solve this problem. However, existing approaches are not generic models for heterogeneous disks in large data centers, e.g, most of the approaches only consider datasets consisting of disks from the same manufacturer (and often of the same disk models). Moreover, some approaches achieve high detection performance in most cases but can not deliver satisfactory results when the datasets of a relatively small amount of disks or have new datasets which have not been seen during training. In this paper, we propose a novel generic disk failure detection approach for heterogeneous disks that can not only deliver a better detective performance but also have good detective adaptability to the disks which have not appeared in training, even when dealing with imbalanced or a relatively small amount of disk datasets. We employ a Long Short-Term Memory (LSTM) based siamese network that can learn the dynamically changed long-term behavior of disk healthy statues. Moreover, this structure can generate a unified and efficient high dimensional disk state embeddings for failure detection of heterogeneous disks. Our evaluation results on two real-world data centers confirm that the proposed system is effective and outperforms several state-of-the-art approaches. Furthermore, we have successfully applied the proposed system to improve the reliability of a data center and exhibit practical long-term availability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Boosting Full-Node Repair in Erasure-Coded StorageShiyao Lin, Guowen Gong, Zhirong Shen, Patrick P. C. Lee 等USENIX ATC 2021 · 被引用 33 次
- Exploit both SMART Attributes and NAND Flash Wear Characteristics to Effectively Forecast SSD-based Storage Failures in ClustersYunfei Gu, Chentao Wu, Xubin HeUSENIX ATC 2024 · 被引用 8 次
- Primo: Practical Learning-Augmented Systems with Interpretable ModelsQinghao Hu, Harsha Nori, Peng Sun, Yonggang Wen 等USENIX ATC 2022 · 被引用 5 次
相关 Paper
- Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced CostJi Zhang, Yuanzhang Wang, Yangtao Wang, Ke Zhou 等DAC 2020 · 被引用 8 次
- NTAM: Neighborhood-Temporal Attention Model for Disk Failure Prediction in Cloud PlatformsChuan Luo, Pu Zhao, Bo Qiao, Youjiang Wu 等WWW 2021 · 被引用 37 次
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 被引用 1,823 次
- Making Disk Failure Predictions SMARTer!Sidi Lu, Bing Luo, Tirthak Patel, Yongtao Yao 等FAST 2020 · 被引用 120 次
- Heterogeneous Anomaly Detection for Software Systems via Semi-supervised Cross-modal AttentionCheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su 等ICSE 2023 · 被引用 52 次
