RL-Watchdog: A Fast and Predictable SSD Liveness Watchdog on Storage Systems
Jinyong Ha, Sangjin Lee, Heon Young Yeom, Yongseok Son
Abstract
This paper proposes a reinforcement learning-based watchdog (RLW) that examines solid-state drive (SSD) liveness or failures by faults (e.g., controller/power faults and high temperature) quickly, precisely, and online to minimize application data loss. To do this, we first provide a lightweight watchdog (LWW) to actively and lightly examine SSD liveness by issuing a liveness-dedicated command to the SSD. Second, we introduce a reinforcement learning-based timeout predictor (RLTP) which predicts the timeout of the dedicated command, enabling the detection of a failure point regardless of the SSD model. Finally, we propose fast failure notification (FFN) to immediately notify the applications of the failure to minimize their potential data loss. We implement RLW with three techniques in a Linux kernel 6.0.0 and evaluate it in a single SSD and RAID using realistic power fault injection. The experimental results reveal that RLW reduces the data loss by up to 96.7% compared with the existing scheme, and its accuracy in predicting failure points reaches up to 99.8%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67f0a32f-c179-4937-b150-ec1e951e16fcBuilds on14
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain et al.NeurIPS 2021 · 149 citations
- A Study of SSD Reliability in Large Scale Enterprise Storage DeploymentsStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2020 · 69 citations
- LeapIO: Efficient and Portable Virtual NVMe Storage on ARM SoCsHuaicheng Li, Mingzhe Hao, Stanko Novakovic, Vaibhav Gogte et al.ASPLOS 2020 · 58 citations
- Hermes: A Fast, Fault-Tolerant and Linearizable Replication ProtocolAntonios Katsarakis, Vasilis Gavrielatos, M. R. Siavash Katebzadeh, Arpit Joshi et al.ASPLOS 2020 · 47 citations
- Operational Characteristics of SSDs in Enterprise Storage Systems: A Large-Scale Field StudyStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2022 · 40 citations
Related papers
- RLAlloc: A Deep Reinforcement Learning-Assisted Resource Allocation Framework for Enhanced Both I/O Throughput and QoS Performance of Multi-Streamed SSDsMengquan Li, Chao Wu, Congming Gao, Cheng Ji et al.DAC 2023 · 3 citations
- Multi-view Feature-based SSD Failure Prediction: What, When, and WhyYuqi Zhang, Wenwen Hao, Ben Niu, Kangkang Liu et al.FAST 2023 · 33 citations
- Reinforcement Learning-Assisted Management for Convertible SSDsQian Wei, Yi Li, Zhiping Jia, Mengying Zhao et al.DAC 2023 · 17 citations
- Reinforcement Learning-Assisted Cache Cleaning to Mitigate Long-Tail Latency in DM-SMRYungang Pan, Zhiping Jia, Zhaoyan Shen, Bingzhe Li et al.DAC 2021 · 12 citations
- Exploit both SMART Attributes and NAND Flash Wear Characteristics to Effectively Forecast SSD-based Storage Failures in ClustersYunfei Gu, Chentao Wu, Xubin HeUSENIX ATC 2024 · 8 citations
