CodeS: Towards Building Open-source Language Models for Text-to-SQL
Haoyang Li, Jing Zhang, Hanbing Liu, Ju Fan, Xiaokang Zhang, Jun Zhu, Renjie Wei, Hongyan Pan, Cuiping Li, Hong Chen
Abstract
Language models have shown promising performance on the task of translating natural language questions into SQL queries (Text-to-SQL). However, most of the state-of-the-art (SOTA) approaches rely on powerful yet closed-source large language models (LLMs), such as ChatGPT and GPT-4, which may have the limitations of unclear model architectures, data privacy risks, and expensive inference overheads. To address the limitations, we introduce CodeS, a series of pre-trained language models with parameters ranging from 1B to 15B, specifically designed for the text-to-SQL task. CodeS is a fully open-source language model, which achieves superior accuracy with much smaller parameter sizes. This paper studies the research challenges in building CodeS. To enhance the SQL generation abilities of CodeS, we adopt an incremental pre-training approach using a specifically curated SQL-centric corpus. Based on this, we address the challenges of schema linking and rapid domain adaptation through strategic prompt construction and a bi-directional data augmentation technique. We conduct comprehensive evaluations on multiple datasets, including the widely used Spider benchmark, the newly released BIRD benchmark, robustness-diagnostic benchmarks such as Spider-DK, Spider-Syn, Spider-Realistic, and Dr.Spider, as well as two real-world datasets created for financial and academic applications. The experimental results show that our CodeS achieves new SOTA accuracy and robustness on nearly all challenging text-to-SQL benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b42da117-53d4-4228-b4d5-7cee03d4449fCited by top-tier papers69
- The Dawn of Natural Language to SQL: Are We Fully Ready? [Experiment, Analysis & Benchmark ]Boyan Li, Yuyu Luo, Chengliang Chai, Guoliang Li et al.VLDB 2024 · 137 citations
- SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement LearningPeixian Ma, Xialie Zhuang, Chengjin Xu, Xuhui Jiang et al.NeurIPS 2025 · 94 citations
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang et al.VLDB 2025 · 90 citations
- Combining Small Language Models and Large Language Models for Zero-Shot NL2SQLJu Fan, Zihui Gu, Songyue Zhang, Yuxin Zhang et al.VLDB 2024 · 71 citations
- MAGIC: Generating Self-Correction Guideline for In-Context Text-to-SQLArian Askari, Christian Pölitz, Xinye TangAAAI 2025 · 44 citations
Builds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
Related papers
- Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL RobustnessShuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan et al.ICLR 2023 · 9 citations
- MCTS-SQL: Light-Weight LLMs Can Master the Text-to-SQL Through Monte Carlo Tree SearchShuozhi Yuan, Liming Chen, Miaomiao Yuan, Jin ZhaoAAAI 2026 · 4 citations
- OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate SupervisionRuilin Hu, Yuyu Luo, Guoliang Li, Shuangqiao Wu et al.VLDB 2026 · 4 citations
- DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link GraphJihyung Lee, Jin-Seop Lee, Jaehoon Lee, YunSeok Choi et al.ACL 2025
- Synthesizing Text-to-SQL Data from Weak and Strong LLMsJiaxi Yang, Binyuan Hui, Min Yang, Jian Yang et al.ACL 2024
