Saving Money for Analytical Workloads in the Cloud
Tapan Srivastava, Raul Castro Fernandez
Abstract
As users migrate their analytical workloads to cloud databases, it is becoming just as important to reduce monetary costs as it is to optimize query runtime. In the cloud, a query is billed based on either its compute time or the amount of data it processes. We observe that analytical queries are either compute- or IO-bound and each query type executes cheaper in a different pricing model. We exploit this opportunity and propose methods to build cheaper execution plans across pricing models that complete within user-defined runtime constraints. We implement these methods and produce execution plans spanning multiple pricing models that reduce the monetary cost for workloads by as much as 56%. We reduce individual query costs by as much as 90%. The prices chosen by cloud vendors for cloud services also impact savings opportunities. To study this effect, we simulate our proposed methods with different cloud prices and observe that multi-cloud savings are robust to changes in cloud vendor prices. These results indicate the massive opportunity to save money by executing workloads across multiple pricing models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d332e1be-70c8-4f43-b672-3a4fae2a57f0Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu et al.VLDB 2020 · 206 citations
- The LDBC Social Network Benchmark: Business Intelligence WorkloadGábor Szárnyas, Jack Waudby, Benjamin A. Steer, Dávid Szakállas et al.VLDB 2023 · 103 citations
- Towards Cost-Optimal Query Processing in the CloudViktor Leis, Maximilian KuschewskiVLDB 2021 · 34 citations
- Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query ProcessingXi Liang, Stavros Sintos, Zechao Shang, Sanjay KrishnanSIGMOD 2021 · 27 citations
- Cloudcast: High-Throughput, Cost-Aware Overlay Multicast in the CloudSarah Wooders, Shu Liu, Paras Jain, Xiangxi Mo et al.NSDI 2024 · 20 citations
Related papers
- Cackle: Analytical Workload Cost and Performance Stability With Elastic PoolsMatthew Perron, Raul Castro Fernandez, David J. DeWitt, Michael J. Cafarella et al.SIGMOD 2024 · 7 citations
- LORE: Learning-Based Resource Recommendation for Big Data QueriesYan Li, Liwei Wang, Bolong Zheng, Zhiyong PengICDE 2025 · 2 citations
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel et al.SIGMOD 2020 · 80 citations
- Riveter: Adaptive Query Suspension and Resumption Framework for Cloud Native DatabasesRui Liu, Aaron J. Elmore, Michael J. Franklin, Sanjay KrishnanICDE 2024 · 1 citation
- AQUOMAN: An Analytic-Query Offloading MachineShuotao Xu, Thomas Bourgeat, Tianhao Huang, Hojun Kim et al.MICRO 2020 · 26 citations
