Multitask Pretraining with Structured Knowledge for Text-to-SQL Generation
Robert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Yang Li, Ming Tan, Parminder Bhatia, Ramesh Nallapati, Xiaofei Ma
摘要
Many machine learning-based low-code or nocode applications involve generating code that interacts with structured knowledge. For example, one of the most studied tasks in this area is generating SQL code from a natural language statement. Prior work shows that incorporating context information from the database schema, such as table and column names, is beneficial to model performance on this task. In this work we present a large pretraining dataset and strategy for learning representations of text, tables, and SQL code that leverages the entire context of the problem. Specifically, we build on existing encoder-decoder architecture by introducing a multitask pretraining framework that complements the unique attributes of our diverse pretraining data. Our work represents the first study on large-scale pretraining of encoderdecoder models for interacting with structured knowledge, and offers a new state-of-the-art foundation model in text-to-SQL generation. We validate our approach with experiments on two SQL tasks, showing improvement over existing methods, including a 1.7 and 2.2 percentage point improvement over prior state-of-thearts on Spider and CoSQL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SQL-Factory: A Multi-Agent Framework for High-Quality and Large-Scale SQL GenerationJiahui Li, Tongwang Wu, Yuren Mao, Yunjun Gao 等VLDB 2026 · 被引用 7 次
- DataVisT5: A Pre-Trained Language Model for Jointly Understanding Text and Data VisualizationZhuoyue Wan, Yuanfeng Song, Shuaimin Li, Chen Jason Zhang 等ICDE 2025 · 被引用 3 次
它引用的顶会 Paper11
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi 等ICLR 2022 · 被引用 347 次
- UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong 等EMNLP 2022 · 被引用 222 次
- Muppet: Massive Multi-task Representations with Pre-FinetuningArmen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen 等EMNLP 2021 · 被引用 176 次
- Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-TrainingPeng Shi, Patrick Ng, Zhiguo Wang, Henghui Zhu 等AAAI 2021 · 被引用 124 次
相关 Paper
- CodeS: Towards Building Open-source Language Models for Text-to-SQLHaoyang Li, Jing Zhang, Hanbing Liu, Ju Fan 等SIGMOD 2024 · 被引用 124 次
- MIGA: A Unified Multi-Task Generation Framework for Conversational Text-to-SQLYingwen Fu, Wenjie Ou, Zhou Yu, Yue LinAAAI 2023 · 被引用 14 次
- MARS-SQL: A Multi-Agent Reinforcement Learning Framework For Text-To-SQLHaolin Yang, Jipeng Zhang, Zhitao He, Alexander Zhou 等ICML 2026 · 被引用 12 次
- MultiTabQA: Generating Tabular Answers for Multi-Table Question AnsweringVaishali Pal, Andrew Yates, Evangelos Kanoulas, Maarten de RijkeACL 2023 · 被引用 7 次
- Few-shot Text-to-SQL Translation using Structure and Content Prompt LearningZihui Gu, Ju Fan, Nan Tang, Lei Cao 等SIGMOD 2023 · 被引用 60 次
