Automating Distributed Tiered Storage Management in Cluster Computing
Herodotos Herodotou, Elena Kakoulli
摘要
Data-intensive platforms such as Hadoop and Spark are routinely used to process massive amounts of data residing on distributed file systems like HDFS. Increasing memory sizes and new hardware technologies (e.g., NVRAM, SSDs) have recently led to the introduction of storage tiering in such settings. However, users are now burdened with the additional complexity of managing the multiple storage tiers and the data residing on them while trying to optimize their workloads. In this paper, we develop a general framework for automatically moving data across the available storage tiers in distributed file systems. Moreover, we employ machine learning for tracking and predicting file access patterns, which we use to decide when and which data to move up or down the storage tiers for increasing system performance. Our approach uses incremental learning to dynamically refine the models with new file accesses, allowing them to naturally adjust and adapt to workload changes over time. Our extensive evaluation using realistic workloads derived from Facebook and CMU traces compares our approach with several other policies and showcases significant benefits in terms of both workload performance and cluster efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Profiling Hyperscale Big Data ProcessingAbraham Gonzalez, Aasheesh Kolli, Samira Manabi Khan, Sihang Liu 等ISCA 2023 · 被引用 30 次
- Proteus: Autonomous Adaptive Storage for Mixed WorkloadsMichael Abebe, Horatiu Lazu, Khuzaima DaudjeeSIGMOD 2022 · 被引用 20 次
- VIP Hashing - Adapting to Skew in Popularity of Data on the FlyAarati Kakaraparthy, Jignesh M. Patel, Brian Kroth, Kwanghyun ParkVLDB 2022 · 被引用 14 次
- HybridTier: an Adaptive and Lightweight CXL-Memory Tiering SystemKevin Song, Jiacheng Yang, Zixuan Wang, Jishen Zhao 等ASPLOS 2025 · 被引用 14 次
- Trident: Task Scheduling over Tiered Storage Systems in Big Data PlatformsHerodotos Herodotou, Elena KakoulliVLDB 2021 · 被引用 10 次
相关 Paper
- MegaMmap: Blurring the Boundary Between Memory and Storage for Data-Intensive WorkloadsLuke Logan, Anthony Kougkas, Xian-He SunSC 2024 · 被引用 3 次
- ArtMem: Adaptive Migration in Reinforcement Learning-Enabled Tiered MemoryXinyue Yi, Hongchao Du, Yu Wang, Jie Zhang 等ISCA 2025 · 被引用 9 次
- ReStore: A Reinforcement Learning Approach for Data Migration in Multi-Tiered StorageTianru Zhang, Tarikul Islam Papon, Teona Bagashvili, Salman Toor 等SIGMOD 2026 · 被引用 1 次
- GMT: GPU Orchestrated Memory Tiering for the Big Data EraChia-Hao Chang, Jihoon Han, Anand Sivasubramaniam, Vikram Sharma Mailthody 等ASPLOS 2024 · 被引用 11 次
- Getting the MOST out of your Storage Hierarchy with Mirror-Optimized Storage TieringKaiwei Tu, Kan Wu, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-DusseauFAST 2026 · 被引用 2 次
