Metasql: A Generate-Then-Rank Framework for Natural Language to SQL Translation
Yuankai Fan, Zhenying He, Tonghui Ren, Can Huang, Yinan Jing, Kai Zhang, X. Sean Wang
Abstract
The Natural Language Interface to Databases (NLIDB) empowers non-technical users with database access through intuitive natural language (NL) interactions. Advanced approaches, utilizing neural sequence-to-sequence models or large-scale language models, typically employ auto-regressive decoding to generate unique SQL queries sequentially. While these translation models have greatly improved the overall translation accuracy, surpassing 70% on NLIDB benchmarks, the use of auto-regressive decoding to generate single SQL queries may result in sub-optimal outputs, potentially leading to erroneous translations. In this paper, we propose Metasql, a unified generate-then-rank framework that can be flexibly incorporated with existing NLIDBs to consistently improve their translation accuracy. Metasql introduces query metadata to control the generation of better SQL query candidates and uses learning-to-rank algorithms to retrieve globally optimized queries. Specifically, Metasql first breaks down the meaning of the given NL query into a set of possible query metadata, representing the basic concepts of the semantics. These metadata are then used as language constraints to steer the underlying translation model toward generating a set of candidate SQL queries. Finally, Metasql ranks the candidates to identify the best matching one for the given NL query. Extensive experiments are performed to study Metasql on two public NLIDB benchmarks. The results show that the performance of the translation models can be effectively improved using Metasql. In particular, applying Metasql to the published Lgesql model obtains a translation accuracy of 77.4 % on the validation set and 72.3 % on the test set of the Spider benchmark, outperforming the baseline by 2.3% and 0.3%, respectively. Moreover, applying Metasql to GpT-4 achieves translation accuracies of 68.6%, 42.0%, and 17.6 % on the three real-world complex scientific databases of Sciencebenchmark, respectively. The code for Metasql is available at https://github.com/Kaimary/MetaSQL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2b6a031-ad50-44d1-b19d-fdb9bf54d009Cited by top-tier papers7
- PURPLE: Making a Large Language Model a Better SQL WriterTonghui Ren, Yuankai Fan, Zhenying He, Ren Huang et al.ICDE 2024 · 49 citations
- Reliable Text-to-SQL with Adaptive AbstentionKaiwen Chen, Yueting Chen, Nick Koudas, Xiaohui YuSIGMOD 2025 · 9 citations
- LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language ModelsWeibin Liao, Xin Gao, Tianyu Jia, Rihong Qiu et al.ICLR 2026 · 9 citations
- The Power of Constraints in Natural Language to SQL TranslationTonghui Ren, Chen Ke, Yuankai Fan, Yinan Jing et al.VLDB 2025 · 4 citations
- Text2VectorSQL: Towards a Unified Interface for Vector Search and SQL QueriesZhengren Wang, Dongwen Yao, Bozhou Li, Dongsheng Ma et al.ICDE 2026 · 1 citation
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 909 citations
- RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQLHaoyang Li, Jing Zhang, Cuiping Li, Hong ChenAAAI 2023 · 343 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Gar: A Generate-and-Rank Approach for Natural Language to SQL TranslationYuankai Fan, Zhenying He, Tonghui Ren, Dianjun Guo et al.ICDE 2023 · 12 citations
- MT-Teql: Evaluating and Augmenting Neural NLIDB on Real-world Linguistic and Schema VariationsPingchuan Ma, Shuai WangVLDB 2022 · 38 citations
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten et al.VLDB 2024 · 65 citations
- Grounding Natural Language to SQL Translation with Data-Based Self-ExplanationsYuankai Fan, Tonghui Ren, Can Huang, Zhenying He et al.ICDE 2025 · 7 citations
- NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL SolutionsShizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta et al.VLDB 2026 · 1 citation
