ROSE: An Intent-Centered Evaluation Metric for NL2SQL
Wenqi Pei, Shizheng Hou, Boyan Li, Chen Han, Zhichao Shi, Yuyu Luo
Abstract
Execution Accuracy (EX), the widely used metric for evaluating the effectiveness of Natural Language to SQL (NL2SQL) solutions, is becoming increasingly unreliable. It is sensitive to syntactic variation, ignores that questions may admit multiple interpretations, and is easily misled by erroneous ground-truth SQL. To address this, we introduce ROSE, an intentcentered metric that focuses on whether the predicted SQL answers the question, rather than consistency with the ground-truth SQL under the reference-dependent paradigm. ROSE employs an adversarial Prover-Refuter cascade: SQL Prover assesses the semantic correctness of a predicted SQL against the user's intent independently, while Adversarial Refuter uses the ground-truth SQL as evidence to challenge and refine this judgment. On our expert-aligned validation set ROSE-VEC, ROSE achieves the best agreement with human experts, outperforming the next-best metric by nearly 24% in Cohen's Kappa. We also conduct a largescale re-evaluation of 19 NL2SQL methods, revealing four valuable insights. We release ROSE and ROSE-VEC to facilitate more reliable NL2SQL research 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6260378b-c5e8-4fc0-91a8-dabe286f85efBuilds on11
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQLHaoyang Li, Jing Zhang, Cuiping Li, Hong ChenAAAI 2023 · 343 citations
- The Dawn of Natural Language to SQL: Are We Fully Ready? [Experiment, Analysis & Benchmark ]Boyan Li, Yuyu Luo, Chengliang Chai, Guoliang Li et al.VLDB 2024 · 137 citations
- CodeS: Towards Building Open-source Language Models for Text-to-SQLHaoyang Li, Jing Zhang, Hanbing Liu, Ju Fan et al.SIGMOD 2024 · 124 citations
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang et al.VLDB 2025 · 90 citations
Related papers
- SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL BenchmarksMohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Oroojlooy, Graham Horwood et al.ACL 2026
- Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL GenerationTarfah Alrashed, Madhup Sukoon, David R. Karger, Natasha F. NoyVLDB 2026
- GBV-SQL: Guided Generation and SQL2Text Back-Translation Validation for Multi-Agent Text2SQLDaojun Chen, Xi Wang, Shenyuan Ren, Qingzhi Ma et al.ACL 2026
- SQLMorph: Query Mutation and Fine-Grained Metrics for Text-to-SQL EvaluationMohammadhossein Malekpour, Mohamed Riahi, Maxime Lamothe, Amine MhedhbiICDE 2026
- Automated Validating and Fixing of Text-to-SQL Translation with Execution ConsistencyYicun Yang, Zhaoguo Wang, Yu Xia, Zhuoran Wei et al.SIGMOD 2025 · 7 citations
