KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers
Chia-Hsuan Lee, Oleksandr Polozov, Matthew Richardson
Abstract
The goal of database question answering is to enable natural language querying of real-life relational databases in diverse application domains. Recently, large-scale datasets such as Spider and WikiSQL facilitated novel modeling techniques for text-to-SQL parsing, improving zero-shot generalization to unseen databases. In this work, we examine the challenges that still prevent these techniques from practical deployment. First, we present KaggleDBQA, a new cross-domain evaluation dataset of real Web databases, with domain-specific data types, original formatting, and unrestricted questions. Second, we re-examine the choice of evaluation tasks for text-to-SQL parsers as applied in real-life settings. Finally, we augment our in-domain evaluation task with database documentation, a naturally occurring source of implicit domain knowledge. We show that KaggleDBQA presents a challenge to state-ofthe-art zero-shot parsers but a more realistic evaluation setting and creative use of associated database documentation boosts their accuracy by over 13.2%, doubling their performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5d21b79-1a87-4182-8c56-a3e6feff43bdCited by top-tier papers31
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- The Dawn of Natural Language to SQL: Are We Fully Ready? [Experiment, Analysis & Benchmark ]Boyan Li, Yuyu Luo, Chengliang Chai, Guoliang Li et al.VLDB 2024 · 137 citations
- Combining Small Language Models and Large Language Models for Zero-Shot NL2SQLJu Fan, Zihui Gu, Songyue Zhang, Yuxin Zhang et al.VLDB 2024 · 71 citations
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten et al.VLDB 2024 · 65 citations
- Is Long Context All You Need? Leveraging LLM's Extended Context for NL2SQLYeounoh Chung, Gaurav Tarlok Kakkar, Yu Gan, Brenton Milne et al.VLDB 2025 · 32 citations
Builds on9
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta et al.AAAI 2020 · 707 citations
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 182 citations
- Exploring Unexplored Generalization Challenges for Cross-Database Semantic ParsingAlane Suhr, Ming-Wei Chang, Peter Shaw, Kenton LeeACL 2020 · 76 citations
- GraPPa: Grammar-Augmented Pre-Training for Table Semantic ParsingTao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang et al.ICLR 2021 · 59 citations
Related papers
- Bridging the Generalization Gap in Text-to-SQL Parsing with Schema ExpansionChen Zhao, Yu Su, Adam Pauls, Emmanouil Antonios PlataniosACL 2022 · 19 citations
- "What Do You Mean by That?" A Parser-Independent Interactive Approach for Enhancing Text-to-SQLYuntao Li, Bei Chen, Qian Liu, Yan Gao et al.EMNLP 2020 · 17 citations
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsFangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao et al.ICLR 2025
- Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL ParsingKun Wu, Lijie Wang, Zhenghua Li, Ao Zhang et al.EMNLP 2021 · 22 citations
- Graphix-T5: Mixing Pre-trained Transformers with Graph-Aware Layers for Text-to-SQL ParsingJinyang Li, Binyuan Hui, Reynold Cheng, Bowen Qin et al.AAAI 2023 · 164 citations
