Sibyl: adaptive and extensible data placement in hybrid storage systems using online reinforcement learning
Gagandeep Singh, Rakesh Nadig, Jisung Park, Rahul Bera, Nastaran Hajinazar, David Novo, Juan Gómez-Luna, Sander Stuijk, Henk Corporaal, Onur Mutlu
Abstract
Hybrid storage systems (HSS) use multiple different storage devices to provide high and scalable storage capacity at high performance. Data placement across different devices is critical to maximize the benefits of such a hybrid system. Recent research proposes various techniques that aim to accurately identify performance-critical data to place it in a "best-fit" storage device. Unfortunately, most of these techniques are rigid, which (1) limits their adaptivity to perform well for a wide range of workloads and storage device configurations, and (2) makes it difficult for designers to extend these techniques to different storage system configurations (e.g., with a different number or different types of storage devices) than the configuration they are designed for. Our goal is to design a new data placement technique for hybrid storage systems that overcomes these issues and provides: (1) adaptivity, by continuously learning from and adapting to the workload and the storage device characteristics, and (2) easy extensibility to a wide range of workloads and HSS configurations.
We introduce Sibyl, the first technique that uses reinforcement learning for data placement in hybrid storage systems. Sibyl observes different features of the running workload as well as the storage devices to make system-aware data placement decisions. For every decision it makes, Sibyl receives a reward from the system that it uses to evaluate the long-term performance impact of its decision and continuously optimizes its data placement policy online.
We implement Sibyl on real systems with various HSS configurations, including dual-and tri-hybrid storage systems, and extensively compare it against four previously proposed data placement techniques (both heuristic-and machine learning-based) over a wide range of workloads. Our results show that Sibyl provides 21.6%/19.9% performance improvement in a performanceoriented/cost-oriented HSS configuration compared to the best previous data placement technique. Our evaluation using an HSS configuration with three different storage devices shows that Sibyl outperforms the state-of-the-art data placement policy by 23.9%-48.2%, while significantly reducing the system architect's burden
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23c0a051-bd5a-4865-a07f-1c79a2a00971Cited by top-tier papers3
- Micro-Armed Bandit: Lightweight & Reusable Reinforcement Learning for Microarchitecture Decision-MakingGerasimos Gerogiannis, Josep TorrellasMICRO 2023 · 23 citations
- IDT: Intelligent Data Placement for Multi-tiered Main Memory with Reinforcement LearningJuneseo Chang, Wanju Doh, Yaebin Moon, Eojin Lee et al.HPDC 2024 · 9 citations
- Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement LearningRahul Bera, Zhenrong Lang, Caroline Hengartner, Konstantinos Kanellopoulos et al.HPCA 2026
Builds on6
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- An Imitation Learning Approach for Cache ReplacementEvan Zheran Liu, Milad Hashemi, Kevin Swersky, Parthasarathy Ranganathan et al.ICML 2020 · 108 citations
- Pythia: A Customizable Hardware Prefetching Framework Using Online Reinforcement LearningRahul Bera, Konstantinos Kanellopoulos, Anant Nori, Taha Shahroodi et al.MICRO 2021 · 95 citations
- Reducing solid-state drive read latency by optimizing read-retryJisung Park, Myungsuk Kim, Myoungjun Chun, Lois Orosa et al.ASPLOS 2021 · 66 citations
- A Deep Reinforcement Learning Framework for Architectural Exploration: A Routerless NoC Case StudyTing-Ru Lin, Drew Penney, Massoud Pedram, Lizhong ChenHPCA 2020 · 48 citations
Related papers
- ReStore: A Reinforcement Learning Approach for Data Migration in Multi-Tiered StorageTianru Zhang, Tarikul Islam Papon, Teona Bagashvili, Salman Toor et al.SIGMOD 2026 · 1 citation
- SAC: A Co-Design Cache Algorithm for Emerging SMR-based High-Density DisksDiansen Sun, Yunpeng ChaiASPLOS 2020 · 13 citations
- ReSemble: Reinforced Ensemble Framework for Data PrefetchingPengmiao Zhang, Rajgopal Kannan, Ajitesh Srivastava, Anant V. Nori et al.SC 2022 · 17 citations
- RLAlloc: A Deep Reinforcement Learning-Assisted Resource Allocation Framework for Enhanced Both I/O Throughput and QoS Performance of Multi-Streamed SSDsMengquan Li, Chao Wu, Congming Gao, Cheng Ji et al.DAC 2023 · 3 citations
- Automating Distributed Tiered Storage Management in Cluster ComputingHerodotos Herodotou, Elena KakoulliVLDB 2020 · 30 citations
