Evaluating Cross-Domain Text-to-SQL Models and Benchmarks
Mohammadreza Pourreza, Davood Rafiei
摘要
Text-to-SQL benchmarks play a crucial role in evaluating the progress made in the field and the ranking of different models. However, accurately matching a model-generated SQL query to a reference SQL query in a benchmark fails for various reasons, such as underspecified natural language queries, inherent assumptions in both model-generated and reference queries, and the non-deterministic nature of SQL output under certain conditions. In this paper, we conduct an extensive study of several prominent cross-domain text-to-SQL benchmarks and re-evaluate some of the top-performing models within these benchmarks, by both manually evaluating the SQL queries and rewriting them in equivalent expressions. Our evaluation reveals that attaining a perfect performance on these benchmarks is unfeasible due to the multiple interpretations that can be derived from the provided samples. Furthermore, we find that the true performance of the models is underestimated and their relative performance changes after a re-evaluation. Most notably, our evaluation reveals a surprising discovery: a recent GPT4-based model surpasses the gold standard reference queries in the Spider benchmark in our human evaluation. This finding highlights the importance of interpreting benchmark evaluations cautiously, while also acknowledging the critical role of additional independent evaluations in driving advancements in the field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Pervasive Annotation Errors Break Text-to-SQL Benchmarks and LeaderboardsTengjun Jin, Yoojin Choi, Yuxuan Zhu, Daniel KangVLDB 2026 · 被引用 7 次
- Dial: A Knowledge-Grounded Dialect-Specific NL2SQL SystemXiang Zhang, Le Zhou, Hongming Xu, Wei Zhou 等VLDB 2026 · 被引用 2 次
- NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL SolutionsShizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta 等VLDB 2026 · 被引用 1 次
- Beyond Static Pipelines: Learning Dynamic Workflows for Text-to-SQLYihan Wang, Peiyu Liu, Runyu Chen, Wei XuICML 2026 · 被引用 1 次
- SQLMorph: Query Mutation and Fine-Grained Metrics for Text-to-SQL EvaluationMohammadhossein Malekpour, Mohamed Riahi, Maxime Lamothe, Amine MhedhbiICDE 2026
它引用的顶会 Paper8
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
- Semantic Evaluation for Text-to-SQL with Distilled Test SuitesRuiqi Zhong, Tao Yu, Dan KleinEMNLP 2020 · 被引用 88 次
- Grounded Adaptation for Zero-shot Executable Semantic ParsingVictor Zhong, Mike Lewis, Sida I. Wang, Luke ZettlemoyerEMNLP 2020 · 被引用 85 次
- Re-examining the Role of Schema Linking in Text-to-SQLWenqiang Lei, Weixin Wang, Zhixin Ma, Tian Gan 等EMNLP 2020 · 被引用 71 次
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov 等ACL 2020 · 被引用 39 次
相关 Paper
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten 等VLDB 2024 · 被引用 65 次
- Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL RobustnessShuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan 等ICLR 2023 · 被引用 9 次
- CodeS: Towards Building Open-source Language Models for Text-to-SQLHaoyang Li, Jing Zhang, Hanbing Liu, Ju Fan 等SIGMOD 2024 · 被引用 124 次
- GBV-SQL: Guided Generation and SQL2Text Back-Translation Validation for Multi-Agent Text2SQLDaojun Chen, Xi Wang, Shenyuan Ren, Qingzhi Ma 等ACL 2026
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsFangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao 等ICLR 2025
