DBPal: A Fully Pluggable NL2SQL Training Pipeline
Nathaniel Weir, Prasetya Ajie Utama, Alex Galakatos, Andrew Crotty, Amir Ilkhechi, Shekar Ramaswamy, Rohin Bhushan, Nadja Geisler, Benjamin Hättasch, Steffen Eger, Ugur Çetintemel, Carsten Binnig
Abstract
Natural language is a promising alternative interface to DBMSs because it enables non-technical users to formulate complex questions in a more concise manner than SQL. Recently, deep learning has gained traction for translating natural language to SQL, since similar ideas have been successful in the related domain of machine translation. However, the core problem with existing deep learning approaches is that they require an enormous amount of training data in order to provide accurate translations. This training data is extremely expensive to curate, since it generally requires humans to manually annotate natural language examples with the corresponding SQL queries (or vice versa).
Based on these observations, we propose DBPal, a new approach that augments existing deep learning techniques in order to improve the performance of models for natural language to SQL translation. More specifically, we present a novel training pipeline that automatically generates synthetic training data in order to (1) improve overall translation accuracy, (2) increase robustness to linguistic variation, and (3) specialize the model for the target database. As we show, our DBPal training pipeline is able to improve both the accuracy and linguistic robustness of state-of-the-art natural language to SQL translation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab22c92c-ac39-45b1-a144-a1540fbc6e0cCited by top-tier papers13
- Constrained Language Models Yield Few-Shot Semantic ParsersRichard Shin, Christopher H. Lin, Sam Thomson, Charles Chen et al.EMNLP 2021 · 131 citations
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang et al.VLDB 2025 · 90 citations
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai et al.SIGMOD 2021 · 90 citations
- CatSQL: Towards Real World Natural Language to SQL ApplicationsHan Fu, Chang Liu, Bin Wu, Feifei Li et al.VLDB 2023 · 79 citations
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten et al.VLDB 2024 · 65 citations
Related papers
- DBTagger: Multi-Task Learning for Keyword Mapping in NLIDBs Using Bi-Directional Recurrent Neural NetworksArif Usta, Akifhan Karakayali, Özgür UlusoyVLDB 2021 · 12 citations
- Natural language to SQL: Where are we today?Hyeonji Kim, Byeong-Hoon So, Wook-Shin Han, Hongrae LeeVLDB 2020 · 148 citations
- Gar: A Generate-and-Rank Approach for Natural Language to SQL TranslationYuankai Fan, Zhenying He, Tonghui Ren, Dianjun Guo et al.ICDE 2023 · 12 citations
- PURPLE: Making a Large Language Model a Better SQL WriterTonghui Ren, Yuankai Fan, Zhenying He, Ren Huang et al.ICDE 2024 · 49 citations
- Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL ParsingKun Wu, Lijie Wang, Zhenghua Li, Ao Zhang et al.EMNLP 2021 · 22 citations
