Facilitating SQL Query Composition and Analysis
Zainab Zolaktaf, Mostafa Milani, Rachel Pottinger
Abstract
Formulating efficient SQL queries requires several cycles of tuning and execution, particularly for inexperienced users. We examine methods that can accelerate and improve this interaction by providing insights about SQL queries prior to execution. We achieve this by predicting properties such as the query answer size, its run-time, and error class. Unlike existing approaches, our approach does not rely on any statistics from the database instance or query execution plans. This is particularly important in settings with limited access to the database instance.
Our approach is based on using data-driven machine learning techniques that rely on large query workloads to model SQL queries and their properties. We evaluate the utility of neural network models and traditional machine learning models. We use two real-world query workloads: the Sloan Digital Sky Survey (SDSS) and the SQLShare query workload. Empirical results show that the neural network models are more accurate in predicting the query error class, achieving a higher F-measure on classes with fewer samples as well as performing better on other problems such as run-time and answer size prediction. These results are encouraging and confirm that SQL query workloads and data-driven machine learning methods can be leveraged to facilitate query composition and analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d093c244-0350-42a8-b77d-41eb7fb4ccfeCited by top-tier papers4
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin et al.SIGMOD 2021 · 113 citations
- Towards Dynamic and Safe Configuration Tuning for Cloud DatabasesXinyi Zhang, Hong Wu, Yang Li, Jian Tan et al.SIGMOD 2022 · 62 citations
- Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query WorkloadJohan Kok Zhi Kang, Gaurav, Sien Yi Tan, Feng Cheng et al.SIGMOD 2021 · 31 citations
- "What makes my queries slow?": Subgroup Discovery for SQL Workload AnalysisYoucef Remil, Anes Bendimerad, Romain Mathonat, Philippe Chaleat et al.ASE 2021 · 12 citations
Related papers
- LORE: Learning-Based Resource Recommendation for Big Data QueriesYan Li, Liwei Wang, Bolong Zheng, Zhiyong PengICDE 2025 · 2 citations
- T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision TreesMaximilian Rieger, Thomas NeumannSIGMOD 2025 · 5 citations
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina et al.VLDB 2020 · 154 citations
- Sibyl: Forecasting Time-Evolving Query WorkloadsHanxian Huang, Tarique Siddiqui, Rana Alotaibi, Carlo Curino et al.SIGMOD 2024 · 14 citations
- A Resource-Aware Deep Cost Model for Big Data Query ProcessingYan Li, Liwei Wang, Sheng Wang, Yuan Sun et al.ICDE 2022 · 13 citations
