OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate Supervision
Ruilin Hu, Yuyu Luo, Guoliang Li, Shuangqiao Wu, Yun Luo
摘要
The Text-to-SQL task enables non-expert users to query structured data through natural language. While recent methods based on closed-source large language models (LLMs) achieve strong performance, their high inference cost, data privacy concerns, and limited transparency hinder real-world deployment. Open-source LLMs are a promising alternative; however, training them for Text-to-SQL remains challenging due to scarce task-specific annotations and the difficulty of learning reliable grounding and reasoning solely from sparse end-to-end supervision. To address these challenges, we present OpenSQL, a data-efficient framework that improves Text-to-SQL performance of open-source LLMs via synthesized intermediate supervision. OpenSQL converts limited (Question, SQL) pairs into rich, task-decomposed training signals that guide the model to learn critical intermediate decisions. Concretely, (1) we train a global-local schema linking module with schema-aware learning to identify and refine relevant tables and columns; (2) we introduce reasoning-enhanced SQL generation, which produces diverse candidates along complementary reasoning paths and selects the best one through stepwise clause-level and semantic-level reasoning; and (3) we design a task-aware data augmentation pipeline that provides the intermediate supervision signals to support the entire training process. With the same 32B LLM backbone, OpenSQL achieves 70.0% accuracy on Bird-dev using only 14 K training samples, outperforming the advanced open-source Text-to-SQL model, OmniSQL, which uses 2.5 M training samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper41
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang 等VLDB 2025 · 被引用 90 次
- Text2sql-Flow: a Robust Sql-Aware Data Augmentation Framework for Text-To-SqlQifeng Cai, Hao Liang, Chang Xu, Tao Xie 等ICDE 2026 · 被引用 1 次
- Alpha-SQL: Zero-Shot Text-to-SQL using Monte Carlo Tree SearchBoyan Li, Jiayi Zhang, Ju Fan, Yanwei Xu 等ICML 2025
- LEAF-SQL: Level-Wise Exploration with Adaptive Fine-Graining for Text-to-SQL Skeleton PredictionZhao Tan, Xiping Liu, Qing Shu, Qizhi Wan 等ICDE 2026
- JOLT-SQL: Joint Loss Tuning of Text-to-SQL with Confusion-aware Noisy Schema SamplingJinwang Song, Hongying Zan, Kunli Zhang, Lingling Mu 等EMNLP 2025
