Tier-Scrubbing: An Adaptive and Tiered Disk Scrubbing Scheme with Improved MTTD and Reduced Cost
Ji Zhang, Yuanzhang Wang, Yangtao Wang, Ke Zhou, Sebastian Schelter, Ping Huang, Bin Cheng, Yongguang Ji
Abstract
Sector errors are a common type of error in modern disks. A sector error that occurs during I/O operations might cause inaccessibility of an application. Even worse, it could result in permanent data loss if the data is being reconstructed, and thereby severely affects the reliability of a storage system. Many disk scrubbing schemes have been proposed to solve this problem. However, existing approaches have several limitations. First, schemes use machine learning (ML) to predict latent sector errors (LSEs), but only leverage a single snapshot of training data to make a prediction, and thereby ignore sequential dependencies between different statuses of a hard disk over time. Second, they accelerate the scrubbing at a fixed rate based on the results of a binary classification model, which may result in unnecessary increases in scrubbing cost. Third, they naively accelerate the scrubbing of the full disk which has LSEs based on the predictive results, but neglect partial high-risk areas (the areas that have a higher probability of encountering LSEs). Lastly, they do not employ strategies to scrub these high-risk areas in advance based on I/O accesses patterns, in order to further increase the efficiency of scrubbing.We address these challenges by designing a Tier-Scrubbing (TS) scheme that combines a Long Short-Term Memory (LSTM) based Adaptive Scrubbing Rate Controller (ASRC), a module focusing on sector error locality to locate high-risk areas in a disk, and a piggyback scrubbing strategy to improve the reliability of a storage system. Our evaluation results on realistic datasets and workloads from two real world data centers demonstrate that TS can simultaneously decrease the Mean-Time-To-Detection (MTTD) by about 80% and the scrubbing cost by 20%, compared to a state-of-the-art scrubbing scheme.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 02097ddb-86a8-47f9-8620-ecc2626a4e69Related papers
- HDDse: Enabling High-Dimensional Disk State Embedding for Generic Failure Detection System of Heterogeneous Disks in Large Data CentersJi Zhang, Ping Huang, Ke Zhou, Ming Xie et al.USENIX ATC 2020 · 20 citations
- DVFS-Based Scrubbing Scheduling for Reliability Maximization on Parallel Tasks in SRAM-based FPGAsRui Li, Heng Yu, Weixiong Jiang, Yajun HaDAC 2020 · 9 citations
- Reinforcement Learning-Assisted Cache Cleaning to Mitigate Long-Tail Latency in DM-SMRYungang Pan, Zhiping Jia, Zhaoyan Shen, Bingzhe Li et al.DAC 2021 · 12 citations
- NTAM: Neighborhood-Temporal Attention Model for Disk Failure Prediction in Cloud PlatformsChuan Luo, Pu Zhao, Bo Qiao, Youjiang Wu et al.WWW 2021 · 37 citations
- ECC Enabled Reliable and Performant Processing-in-MemoryJeageun Jung, Margaret Lee, Mattan ErezISCA 2026
