Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query Workload
Johan Kok Zhi Kang, Gaurav, Sien Yi Tan, Feng Cheng, Shixuan Sun, Bingsheng He
摘要
The use of deep learning models for forecasting the resource consumption patterns of SQL queries have recently been a popular area of study. With many companies using cloud platforms to power their data lakes for large scale analytic demands, these models form a critical part of the pipeline in managing cloud resource provisioning. While these models have demonstrated promising accuracy, training them over large scale industry workloads are expensive. Space inefficiencies of encoding techniques over large numbers of queries and excessive padding used to enforce shape consistency across diverse query plans implies 1) longer model training time and 2) the need for expensive, scaled up infrastructure to support batched training. In turn, we developed Prestroid, a tree convolution based data science pipeline that accurately predicts resource consumption patterns of query traces, but at a much lower cost. We evaluated our pipeline over 19K Presto OLAP queries from Grab, on a data lake of more than 20PB of data. Experimental results imply that our pipeline outperforms benchmarks on predictive accuracy, contributing to more precise resource prediction for large-scale workloads, yet also reduces per-batch memory footprint by 13.5x and per-epoch training time by 3.45x. We demonstrate direct cost savings of up to 13.2x for large batched model training over Microsoft Azure VMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Zero-Shot Cost Models for Out-of-the-box Learned Cost PredictionBenjamin Hilprecht, Carsten BinnigVLDB 2022 · 被引用 90 次
- LIO: A lightweight and interpretable query optimizer based on an evolutionary forestChen Ye, Shujie Ma, Guojun Dai, Hengtong ZhangVLDB 2026 · 被引用 1 次
- Micro: a Lightweight Middleware for Optimizing Cross-Store Cross-Model Graph-Relation JoinsXiuwen Zheng, Arun Kumar, Amarnath GuptaICDE 2026
- AXE: A Task Decomposition Approach to Learned LSM TuningAndy Huynh, Anwesha Saha, Harshal A. Chaudhari, Manos AthanassoulisVLDB 2025
它引用的顶会 Paper5
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul 等SIGMOD 2021 · 被引用 242 次
- Efficient Segmentation: Learning Downsampling Near Semantic BoundariesDmitrii Marin, Zijian He, Peter Vajda, Priyam Chatterjee 等ICCV 2019 · 被引用 107 次
- QuickSel: Quick Selectivity Learning with Mixture ModelsYongjoo Park, Shucheng Zhong, Barzan MozafariSIGMOD 2020 · 被引用 66 次
- Facilitating SQL Query Composition and AnalysisZainab Zolaktaf, Mostafa Milani, Rachel PottingerSIGMOD 2020 · 被引用 19 次
相关 Paper
- A Resource-Aware Deep Cost Model for Big Data Query ProcessingYan Li, Liwei Wang, Sheng Wang, Yuan Sun 等ICDE 2022 · 被引用 13 次
- LORE: Learning-Based Resource Recommendation for Big Data QueriesYan Li, Liwei Wang, Bolong Zheng, Zhiyong PengICDE 2025 · 被引用 2 次
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel 等SIGMOD 2020 · 被引用 80 次
- End-to-end Optimization of Machine Learning Prediction QueriesKwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen 等SIGMOD 2022 · 被引用 50 次
- Understanding and Detecting Query Performance Regression in Practical Index Tuning: [Experiments & Analysis]Wentao Wu, Anshuman Dutt, Gaoxiang Xu, Vivek R. Narasayya 等SIGMOD 2026
