Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query Workload
Johan Kok Zhi Kang, Gaurav, Sien Yi Tan, Feng Cheng, Shixuan Sun, Bingsheng He
Abstract
The use of deep learning models for forecasting the resource consumption patterns of SQL queries have recently been a popular area of study. With many companies using cloud platforms to power their data lakes for large scale analytic demands, these models form a critical part of the pipeline in managing cloud resource provisioning. While these models have demonstrated promising accuracy, training them over large scale industry workloads are expensive. Space inefficiencies of encoding techniques over large numbers of queries and excessive padding used to enforce shape consistency across diverse query plans implies 1) longer model training time and 2) the need for expensive, scaled up infrastructure to support batched training. In turn, we developed Prestroid, a tree convolution based data science pipeline that accurately predicts resource consumption patterns of query traces, but at a much lower cost. We evaluated our pipeline over 19K Presto OLAP queries from Grab, on a data lake of more than 20PB of data. Experimental results imply that our pipeline outperforms benchmarks on predictive accuracy, contributing to more precise resource prediction for large-scale workloads, yet also reduces per-batch memory footprint by 13.5x and per-epoch training time by 3.45x. We demonstrate direct cost savings of up to 13.2x for large batched model training over Microsoft Azure VMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2ab74bc-7300-43ca-9c3c-c70812fc3cf6Cited by top-tier papers4
- Zero-Shot Cost Models for Out-of-the-box Learned Cost PredictionBenjamin Hilprecht, Carsten BinnigVLDB 2022 · 90 citations
- LIO: A lightweight and interpretable query optimizer based on an evolutionary forestChen Ye, Shujie Ma, Guojun Dai, Hengtong ZhangVLDB 2026 · 1 citation
- Micro: a Lightweight Middleware for Optimizing Cross-Store Cross-Model Graph-Relation JoinsXiuwen Zheng, Arun Kumar, Amarnath GuptaICDE 2026
- AXE: A Task Decomposition Approach to Learned LSM TuningAndy Huynh, Anwesha Saha, Harshal A. Chaudhari, Manos AthanassoulisVLDB 2025
Builds on5
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 251 citations
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul et al.SIGMOD 2021 · 242 citations
- Efficient Segmentation: Learning Downsampling Near Semantic BoundariesDmitrii Marin, Zijian He, Peter Vajda, Priyam Chatterjee et al.ICCV 2019 · 107 citations
- QuickSel: Quick Selectivity Learning with Mixture ModelsYongjoo Park, Shucheng Zhong, Barzan MozafariSIGMOD 2020 · 66 citations
- Facilitating SQL Query Composition and AnalysisZainab Zolaktaf, Mostafa Milani, Rachel PottingerSIGMOD 2020 · 19 citations
Related papers
- A Resource-Aware Deep Cost Model for Big Data Query ProcessingYan Li, Liwei Wang, Sheng Wang, Yuan Sun et al.ICDE 2022 · 13 citations
- LORE: Learning-Based Resource Recommendation for Big Data QueriesYan Li, Liwei Wang, Bolong Zheng, Zhiyong PengICDE 2025 · 2 citations
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel et al.SIGMOD 2020 · 80 citations
- End-to-end Optimization of Machine Learning Prediction QueriesKwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen et al.SIGMOD 2022 · 50 citations
- Understanding and Detecting Query Performance Regression in Practical Index Tuning: [Experiments & Analysis]Wentao Wu, Anshuman Dutt, Gaoxiang Xu, Vivek R. Narasayya et al.SIGMOD 2026
