Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL Generation
Tarfah Alrashed, Madhup Sukoon, David R. Karger, Natasha F. Noy
摘要
Large Language Models (LLMs) have achieved impressive performance in translating natural language queries into executable SQL. However, these systems remain prone to deceptive failures : generating syntactically valid SQL that executes but fails to capture the user's intent. In this work, we argue that further progress can be made by focusing specifically on verification. In this paper, we formalize Text-to-SQL verification as a standalone task: assessing whether a candidate SQL query semantically satisfies a natural language request, without access to ground-truth labels. We propose and compare two modular verification strategies: Round-Trip Critique , which reverse-translates SQL into natural language to detect semantic drift, and Synthetic Execution Consistency , which uses unit-test-like synthetic inputs to ground verification in execution results. Our evaluation shows that these methods provide a robust signal for identifying incorrect queries, successfully flagging 64% of errors from a state-of-the-art generator and outperforming standard error detectors. We demonstrate two critical applications: (1) auditing foundational benchmarks (Spider, BIRD and KaggleDBQA), revealing that over two-thirds of "generator failures" actually stem from flawed benchmark labels, and (2) enabling Selective Generation, where a system uses verification signals to abstain from answering when confidence is low. Our results show that this paradigm significantly improves the quality of deployed data interfaces by transforming silent failures into explicit abstentions that alert the user to the failure and give them a chance to correct it.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun 等VLDB 2024 · 被引用 609 次
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and InconsistenciesSunnie S. Y. Kim, Jennifer Wortman Vaughan, Q. Vera Liao, Tania Lombrozo 等CHI 2025 · 被引用 118 次
相关 Paper
- GBV-SQL: Guided Generation and SQL2Text Back-Translation Validation for Multi-Agent Text2SQLDaojun Chen, Xi Wang, Shenyuan Ren, Qingzhi Ma 等ACL 2026
- SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQLYue Gong, Chuan Lei, Xiao Qin, Kapil Vaidya 等NeurIPS 2025 · 被引用 21 次
- SafeQL: Search-based Refinement for Safe and Efficient LLM-based Text-to-SQLGeonho Lee, Min-Soo KimVLDB 2026
- SQL-Checker: Error Detection and Labeling for Text-to-SQL with Interpretability AnalysisXingyu Ma, Xin Tian, Lingxiang Wu, Xuepeng Wang 等WWW 2026
- SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL BenchmarksMohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Oroojlooy, Graham Horwood 等ACL 2026
