USENIX ATC2024顶会
RL-Watchdog: A Fast and Predictable SSD Liveness Watchdog on Storage Systems
Jinyong Ha, Sangjin Lee, Heon Young Yeom, Yongseok Son
摘要
This paper proposes a reinforcement learning-based watchdog (RLW) that examines solid-state drive (SSD) liveness or failures by faults (e.g., controller/power faults and high temperature) quickly, precisely, and online to minimize application data loss. To do this, we first provide a lightweight watchdog (LWW) to actively and lightly examine SSD liveness by issuing a liveness-dedicated command to the SSD. Second, we introduce a reinforcement learning-based timeout predictor (RLTP) which predicts the timeout of the dedicated command, enabling the detection of a failure point regardless of the SSD model. Finally, we propose fast failure notification (FFN) to immediately notify the applications of the failure to minimize their potential data loss. We implement RLW with three techniques in a Linux kernel 6.0.0 and evaluate it in a single SSD and RAID using realistic power fault injection. The experimental results reveal that RLW reduces the data loss by up to 96.7% compared with the existing scheme, and its accuracy in predicting failure points reaches up to 99.8%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain 等NeurIPS 2021 · 被引用 149 次
- A Study of SSD Reliability in Large Scale Enterprise Storage DeploymentsStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2020 · 被引用 69 次
- LeapIO: Efficient and Portable Virtual NVMe Storage on ARM SoCsHuaicheng Li, Mingzhe Hao, Stanko Novakovic, Vaibhav Gogte 等ASPLOS 2020 · 被引用 58 次
- Hermes: A Fast, Fault-Tolerant and Linearizable Replication ProtocolAntonios Katsarakis, Vasilis Gavrielatos, M. R. Siavash Katebzadeh, Arpit Joshi 等ASPLOS 2020 · 被引用 47 次
- Operational Characteristics of SSDs in Enterprise Storage Systems: A Large-Scale Field StudyStathis Maneas, Kaveh Mahdaviani, Tim Emami, Bianca SchroederFAST 2022 · 被引用 40 次
相关 Paper
- RLAlloc: A Deep Reinforcement Learning-Assisted Resource Allocation Framework for Enhanced Both I/O Throughput and QoS Performance of Multi-Streamed SSDsMengquan Li, Chao Wu, Congming Gao, Cheng Ji 等DAC 2023 · 被引用 3 次
- Multi-view Feature-based SSD Failure Prediction: What, When, and WhyYuqi Zhang, Wenwen Hao, Ben Niu, Kangkang Liu 等FAST 2023 · 被引用 33 次
- Reinforcement Learning-Assisted Management for Convertible SSDsQian Wei, Yi Li, Zhiping Jia, Mengying Zhao 等DAC 2023 · 被引用 17 次
- Reinforcement Learning-Assisted Cache Cleaning to Mitigate Long-Tail Latency in DM-SMRYungang Pan, Zhiping Jia, Zhaoyan Shen, Bingzhe Li 等DAC 2021 · 被引用 12 次
- Exploit both SMART Attributes and NAND Flash Wear Characteristics to Effectively Forecast SSD-based Storage Failures in ClustersYunfei Gu, Chentao Wu, Xubin HeUSENIX ATC 2024 · 被引用 8 次
