Automating Distributed Tiered Storage Management in Cluster Computing
Herodotos Herodotou, Elena Kakoulli
Abstract
Data-intensive platforms such as Hadoop and Spark are routinely used to process massive amounts of data residing on distributed file systems like HDFS. Increasing memory sizes and new hardware technologies (e.g., NVRAM, SSDs) have recently led to the introduction of storage tiering in such settings. However, users are now burdened with the additional complexity of managing the multiple storage tiers and the data residing on them while trying to optimize their workloads. In this paper, we develop a general framework for automatically moving data across the available storage tiers in distributed file systems. Moreover, we employ machine learning for tracking and predicting file access patterns, which we use to decide when and which data to move up or down the storage tiers for increasing system performance. Our approach uses incremental learning to dynamically refine the models with new file accesses, allowing them to naturally adjust and adapt to workload changes over time. Our extensive evaluation using realistic workloads derived from Facebook and CMU traces compares our approach with several other policies and showcases significant benefits in terms of both workload performance and cluster efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5249493e-97fe-40af-8967-44ea30a1408bCited by top-tier papers7
- Profiling Hyperscale Big Data ProcessingAbraham Gonzalez, Aasheesh Kolli, Samira Manabi Khan, Sihang Liu et al.ISCA 2023 · 30 citations
- Proteus: Autonomous Adaptive Storage for Mixed WorkloadsMichael Abebe, Horatiu Lazu, Khuzaima DaudjeeSIGMOD 2022 · 20 citations
- VIP Hashing - Adapting to Skew in Popularity of Data on the FlyAarati Kakaraparthy, Jignesh M. Patel, Brian Kroth, Kwanghyun ParkVLDB 2022 · 14 citations
- HybridTier: an Adaptive and Lightweight CXL-Memory Tiering SystemKevin Song, Jiacheng Yang, Zixuan Wang, Jishen Zhao et al.ASPLOS 2025 · 14 citations
- Trident: Task Scheduling over Tiered Storage Systems in Big Data PlatformsHerodotos Herodotou, Elena KakoulliVLDB 2021 · 10 citations
Related papers
- MegaMmap: Blurring the Boundary Between Memory and Storage for Data-Intensive WorkloadsLuke Logan, Anthony Kougkas, Xian-He SunSC 2024 · 3 citations
- ArtMem: Adaptive Migration in Reinforcement Learning-Enabled Tiered MemoryXinyue Yi, Hongchao Du, Yu Wang, Jie Zhang et al.ISCA 2025 · 9 citations
- ReStore: A Reinforcement Learning Approach for Data Migration in Multi-Tiered StorageTianru Zhang, Tarikul Islam Papon, Teona Bagashvili, Salman Toor et al.SIGMOD 2026 · 1 citation
- GMT: GPU Orchestrated Memory Tiering for the Big Data EraChia-Hao Chang, Jihoon Han, Anand Sivasubramaniam, Vikram Sharma Mailthody et al.ASPLOS 2024 · 11 citations
- Getting the MOST out of your Storage Hierarchy with Mirror-Optimized Storage TieringKaiwei Tu, Kan Wu, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-DusseauFAST 2026 · 2 citations
