Evaluating Cross-Domain Text-to-SQL Models and Benchmarks
Mohammadreza Pourreza, Davood Rafiei
Abstract
Text-to-SQL benchmarks play a crucial role in evaluating the progress made in the field and the ranking of different models. However, accurately matching a model-generated SQL query to a reference SQL query in a benchmark fails for various reasons, such as underspecified natural language queries, inherent assumptions in both model-generated and reference queries, and the non-deterministic nature of SQL output under certain conditions. In this paper, we conduct an extensive study of several prominent cross-domain text-to-SQL benchmarks and re-evaluate some of the top-performing models within these benchmarks, by both manually evaluating the SQL queries and rewriting them in equivalent expressions. Our evaluation reveals that attaining a perfect performance on these benchmarks is unfeasible due to the multiple interpretations that can be derived from the provided samples. Furthermore, we find that the true performance of the models is underestimated and their relative performance changes after a re-evaluation. Most notably, our evaluation reveals a surprising discovery: a recent GPT4-based model surpasses the gold standard reference queries in the Spider benchmark in our human evaluation. This finding highlights the importance of interpreting benchmark evaluations cautiously, while also acknowledging the critical role of additional independent evaluations in driving advancements in the field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9fdd49a-0713-46d6-b320-2169e3cf8f92Cited by top-tier papers5
- Pervasive Annotation Errors Break Text-to-SQL Benchmarks and LeaderboardsTengjun Jin, Yoojin Choi, Yuxuan Zhu, Daniel KangVLDB 2026 · 7 citations
- Dial: A Knowledge-Grounded Dialect-Specific NL2SQL SystemXiang Zhang, Le Zhou, Hongming Xu, Wei Zhou et al.VLDB 2026 · 2 citations
- NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL SolutionsShizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta et al.VLDB 2026 · 1 citation
- Beyond Static Pipelines: Learning Dynamic Workflows for Text-to-SQLYihan Wang, Peiyu Liu, Runyu Chen, Wei XuICML 2026 · 1 citation
- SQLMorph: Query Mutation and Fine-Grained Metrics for Text-to-SQL EvaluationMohammadhossein Malekpour, Mohamed Riahi, Maxime Lamothe, Amine MhedhbiICDE 2026
Builds on8
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 909 citations
- Semantic Evaluation for Text-to-SQL with Distilled Test SuitesRuiqi Zhong, Tao Yu, Dan KleinEMNLP 2020 · 88 citations
- Grounded Adaptation for Zero-shot Executable Semantic ParsingVictor Zhong, Mike Lewis, Sida I. Wang, Luke ZettlemoyerEMNLP 2020 · 85 citations
- Re-examining the Role of Schema Linking in Text-to-SQLWenqiang Lei, Weixin Wang, Zhixin Ma, Tian Gan et al.EMNLP 2020 · 71 citations
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov et al.ACL 2020 · 39 citations
Related papers
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten et al.VLDB 2024 · 65 citations
- Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL RobustnessShuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan et al.ICLR 2023 · 9 citations
- CodeS: Towards Building Open-source Language Models for Text-to-SQLHaoyang Li, Jing Zhang, Hanbing Liu, Ju Fan et al.SIGMOD 2024 · 124 citations
- GBV-SQL: Guided Generation and SQL2Text Back-Translation Validation for Multi-Agent Text2SQLDaojun Chen, Xi Wang, Shenyuan Ren, Qingzhi Ma et al.ACL 2026
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsFangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao et al.ICLR 2025
