Towards Optimizing Storage Costs on the Cloud
Koyel Mukherjee, Raunak Shah, Shiv Kumar Saini, Karanpreet Singh, Khushi, Harsh Kesarwani, Kavya Barnwal, Ayush Chauhan
摘要
We study the problem of optimizing data storage and access costs on the cloud while ensuring that the desired performance or latency is unaffected. We first propose an optimizer that optimizes the data placement tier (on the cloud) and the choice of compression schemes to apply, for given data partitions with temporal access predictions. Secondly, we propose a model to learn the compression performance of multiple algorithms across data partitions in different formats to generate compression performance predictions on the fly, as inputs to the optimizer. Thirdly, we propose to approach the data partitioning problem fundamentally differently than the current default in most data lakes where partitioning is in the form of ingestion batches. We propose access pattern aware data partitioning and formulate an optimization problem that optimizes the size and reading costs of partitions subject to access patterns.
We study the various optimization problems theoretically as well as empirically, and provide theoretical bounds as well as hardness results. We propose a unified pipeline of cost minimization, called SCOPe that combines the different modules. We extensively compare the performance of our methods with related baselines from the literature on TPC-H data as well as enterprise datasets (ranging from GB to PB in volume) and show that SCOPe substantially improves over the baselines. We show significant cost savings compared to platform baselines, of the order of 50% to 83% on enterprise Data Lake datasets that range from terabytes to petabytes in volume.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Partition, Don't Sort! Compression Boosters for Cloud Data Ingestion PipelinesPatrick Hansert, Sebastian MichelVLDB 2024 · 被引用 2 次
- TCO-driven Storage Provisioning for Exascale Data CentersTimothy Kim, Saurabh Kadekodi, Arif Merchant, Prashant Nema 等EuroSys 2026
它引用的顶会 Paper6
- Quiver: An Informed Storage Cache for Deep LearningAbhishek Vijaya Kumar, Muthian SivathanuFAST 2020 · 被引用 91 次
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu 等VLDB 2021 · 被引用 67 次
- CompressDB: Enabling Efficient Compressed Data Direct Processing for Various DatabasesFeng Zhang, Weitao Wan, Chenyang Zhang, Jidong Zhai 等SIGMOD 2022 · 被引用 46 次
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang 等SIGMOD 2021 · 被引用 37 次
- Cost Modelling for Optimal Data Placement in Heterogeneous Main MemoryRobert Lasch, Thomas Legler, Norman May, Bernhard Scheirle 等VLDB 2022 · 被引用 12 次
相关 Paper
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel 等SIGMOD 2020 · 被引用 80 次
- SkyPIE: A Fast & Accurate Oracle for Object PlacementTiemo Bang, Chris Douglas, Natacha Crooks, Joseph M. HellersteinSIGMOD 2024
- Phoebe: A Learning-based Checkpoint OptimizerYiwen Zhu, Matteo Interlandi, Abhishek Roy, Krishnadhan Das 等VLDB 2021 · 被引用 10 次
- PAW: Data Partitioning Meets Workload VarianceZhe Li, Man Lung Yiu, Tsz Nam ChanICDE 2022 · 被引用 6 次
- Budget-Conscious Fine-Grained Configuration Optimization for Spatio-Temporal ApplicationsKeven Richly, Rainer Schlosser, Martin BoissierVLDB 2022 · 被引用 3 次
