Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness
Shuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang, Jiarong Jiang, Joseph Lilien, Steve Ash, William Yang Wang
Abstract
Neural text-to-SQL models have achieved remarkable performance in translating natural language questions into SQL queries. However, recent studies reveal that text-to-SQL models are vulnerable to task-specific perturbations. Previous curated robustness test sets usually focus on individual phenomena. In this paper, we propose a comprehensive robustness benchmark 1 based on Spider, a cross-domain text-to-SQL benchmark, to diagnose the model robustness. We design 17 perturbations on databases, natural language questions, and SQL queries to measure the robustness from different angles. In order to collect more diversified natural question perturbations, we utilize large pretrained language models (PLMs) to simulate human behaviors in creating natural questions. We conduct a diagnostic study of the state-of-the-art models on the robustness set. Experimental results reveal that even the most robust model suffers from a 14.0% performance drop overall and a 50.7% performance drop on the most challenging perturbation. We also present a breakdown analysis regarding text-to-SQL model designs and provide insights for improving model robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 201af16c-e022-4521-af4e-6690884c48b9Cited by top-tier papers6
- The Dawn of Natural Language to SQL: Are We Fully Ready? [Experiment, Analysis & Benchmark ]Boyan Li, Yuyu Luo, Chengliang Chai, Guoliang Li et al.VLDB 2024 · 137 citations
- RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial PerturbationsYilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi et al.ACL 2023 · 7 citations
- DIVER: A Robust Text-to-SQL System with Dynamic Interactive Value Linking and Evidence ReasoningYafeng Nan, Haifeng Sun, Zirui Zhuang, Qi Qi et al.SIGMOD 2026 · 3 citations
- TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database QueriesChao Deng, Ju Fan, Yuyu Luo, Qinliang Xue et al.VLDB 2026 · 1 citation
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQLYou Peng, Youhe Jiang, Wenqi Jiang, Chen Wang et al.ICDE 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 340 citations
Related papers
- Towards Robustness of Text-to-SQL Models against Synonym SubstitutionYujian Gan, Xinyun Chen, Qiuping Huang, Matthew Purver et al.ACL 2021
- SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL BenchmarksMohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Oroojlooy, Graham Horwood et al.ACL 2026
- EVOSCHEMA: TOWARDS TEXT-TO-SQL ROBUSTNESS AGAINST SCHEMA EVOLUTIONTianshu Zhang, Kun Qian, Siddhartha Sahai, Yuan Tian et al.VLDB 2025 · 3 citations
- Evaluating Cross-Domain Text-to-SQL Models and BenchmarksMohammadreza Pourreza, Davood RafieiEMNLP 2023 · 14 citations
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten et al.VLDB 2024 · 65 citations
