Multitask Pretraining with Structured Knowledge for Text-to-SQL Generation
Robert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Yang Li, Ming Tan, Parminder Bhatia, Ramesh Nallapati, Xiaofei Ma
Abstract
Many machine learning-based low-code or nocode applications involve generating code that interacts with structured knowledge. For example, one of the most studied tasks in this area is generating SQL code from a natural language statement. Prior work shows that incorporating context information from the database schema, such as table and column names, is beneficial to model performance on this task. In this work we present a large pretraining dataset and strategy for learning representations of text, tables, and SQL code that leverages the entire context of the problem. Specifically, we build on existing encoder-decoder architecture by introducing a multitask pretraining framework that complements the unique attributes of our diverse pretraining data. Our work represents the first study on large-scale pretraining of encoderdecoder models for interacting with structured knowledge, and offers a new state-of-the-art foundation model in text-to-SQL generation. We validate our approach with experiments on two SQL tasks, showing improvement over existing methods, including a 1.7 and 2.2 percentage point improvement over prior state-of-thearts on Spider and CoSQL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da6af802-496d-428d-bcf3-21c2ee0e2c5dCited by top-tier papers2
- SQL-Factory: A Multi-Agent Framework for High-Quality and Large-Scale SQL GenerationJiahui Li, Tongwang Wu, Yuren Mao, Yunjun Gao et al.VLDB 2026 · 7 citations
- DataVisT5: A Pre-Trained Language Model for Jointly Understanding Text and Data VisualizationZhuoyue Wan, Yuanfeng Song, Shuaimin Li, Chen Jason Zhang et al.ICDE 2025 · 3 citations
Builds on11
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong et al.EMNLP 2022 · 222 citations
- Muppet: Massive Multi-task Representations with Pre-FinetuningArmen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen et al.EMNLP 2021 · 176 citations
- Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-TrainingPeng Shi, Patrick Ng, Zhiguo Wang, Henghui Zhu et al.AAAI 2021 · 124 citations
Related papers
- CodeS: Towards Building Open-source Language Models for Text-to-SQLHaoyang Li, Jing Zhang, Hanbing Liu, Ju Fan et al.SIGMOD 2024 · 124 citations
- MIGA: A Unified Multi-Task Generation Framework for Conversational Text-to-SQLYingwen Fu, Wenjie Ou, Zhou Yu, Yue LinAAAI 2023 · 14 citations
- MARS-SQL: A Multi-Agent Reinforcement Learning Framework For Text-To-SQLHaolin Yang, Jipeng Zhang, Zhitao He, Alexander Zhou et al.ICML 2026 · 12 citations
- MultiTabQA: Generating Tabular Answers for Multi-Table Question AnsweringVaishali Pal, Andrew Yates, Evangelos Kanoulas, Maarten de RijkeACL 2023 · 7 citations
- Few-shot Text-to-SQL Translation using Structure and Content Prompt LearningZihui Gu, Ju Fan, Nan Tang, Lei Cao et al.SIGMOD 2023 · 60 citations
