ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL Systems
Yi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten, Georgia Koutrika, Kurt Stockinger
摘要
Natural Language to SQL systems (NL-to-SQL) have recently shown improved accuracy (exceeding 80%) for natural language to SQL query translation due to the emergence of transformer-based language models, and the popularity of the Spider benchmark. However, Spider mainly contains simple databases with few tables, columns, and entries, which do not reflect a realistic setting. Moreover, complex real-world databases with domain-specific content have little to no training data available in the form of NL/SQL-pairs leading to poor performance of existing NL-to-SQL systems.
In this paper, we introduce ScienceBenchmark , a new complex NL-to-SQL benchmark for three real-world, highly domain-specific databases. For this new benchmark, SQL experts and domain experts created high-quality NL/SQL-pairs for each domain. To garner more data, we extended the small amount of human-generated data with synthetic data generated using GPT-3. We show that our benchmark is highly challenging, as the top performing systems on Spider achieve a very low performance on our benchmark. Thus, the challenge is many-fold: creating NL-to-SQL systems for highly complex domains with a small amount of hand-made training data augmented with synthetic data. To our knowledge, ScienceBenchmark is the first NL-to-SQL benchmark designed with complex real-world scientific databases, containing challenging training and test data carefully validated by domain experts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- CodeS: Towards Building Open-source Language Models for Text-to-SQLHaoyang Li, Jing Zhang, Hanbing Liu, Ju Fan 等SIGMOD 2024 · 被引用 124 次
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang 等VLDB 2025 · 被引用 90 次
- Reward-SQL: Boosting Text-to-SQL via Stepwise Execution-Aware Reasoning and Process-Supervised RewardsYuxin Zhang, Meihao Fan, Ju Fan, Mingyang Yi 等SIGMOD 2026 · 被引用 24 次
- Metasql: A Generate-Then-Rank Framework for Natural Language to SQL TranslationYuankai Fan, Zhenying He, Tonghui Ren, Can Huang 等ICDE 2024 · 被引用 23 次
- Sphinteract: Resolving Ambiguities in NL2SQL Through User InteractionFuheng Zhao, Shaleen Deep, Fotis Psallidas, Avrilia Floratou 等VLDB 2025 · 被引用 12 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQLHaoyang Li, Jing Zhang, Cuiping Li, Hong ChenAAAI 2023 · 被引用 343 次
- Unnatural Instructions: Tuning Language Models with (Almost) No Human LaborOr Honovich, Thomas Scialom, Omer Levy, Timo SchickACL 2023 · 被引用 92 次
- GraPPa: Grammar-Augmented Pre-Training for Table Semantic ParsingTao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang 等ICLR 2021 · 被引用 59 次
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov 等ACL 2020 · 被引用 39 次
相关 Paper
- Evaluating Cross-Domain Text-to-SQL Models and BenchmarksMohammadreza Pourreza, Davood RafieiEMNLP 2023 · 被引用 14 次
- Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL RobustnessShuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan 等ICLR 2023 · 被引用 9 次
- NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL SolutionsShizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta 等VLDB 2026 · 被引用 1 次
- TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database QueriesChao Deng, Ju Fan, Yuyu Luo, Qinliang Xue 等VLDB 2026 · 被引用 1 次
- SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL BenchmarksMohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Oroojlooy, Graham Horwood 等ACL 2026
