Synthesizing Text-to-SQL Data from Weak and Strong LLMs
Jiaxi Yang, Binyuan Hui, Min Yang, Jian Yang, Junyang Lin, Chang Zhou
摘要
The capability gap between open-source and closed-source large language models (LLMs) remains challenging in text-to-SQL tasks. In this paper, we introduce a synthetic data approach that amalgamates strong data generated by larger, more potent models (strong models) with weak data produced by smaller, less wellaligned models (weak models). Our approach contributes to the improvement of domain generalization in text-to-SQL models and investigates the potential of weak data supervision through preference learning. Moreover, we utilize the synthetic data approach for instruction tuning on open-source LLMs, yielding SENSE, a specialized text-to-SQL model. The effectiveness of SENSE is substantiated by achieving state-of-the-art results on the SPIDER and BIRD benchmarks, thereby mitigating the performance disparity between open-source models and the methods derived from closed-source models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement LearningPeixian Ma, Xialie Zhuang, Chengjin Xu, Xuhui Jiang 等NeurIPS 2025 · 被引用 94 次
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang 等VLDB 2025 · 被引用 90 次
- DeepEye-SQL: A Software-Engineering-Inspired Text-to-SQL FrameworkBoyan Li, Chong Chen, Zhujun Xue, Yinan Mei 等SIGMOD 2026 · 被引用 40 次
- Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQLHanbing Liu, Haoyang Li, Xiaokang Zhang, Ruotong Chen 等ACL 2025 · 被引用 10 次
- LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language ModelsWeibin Liao, Xin Gao, Tianyu Jia, Rihong Qiu 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
相关 Paper
- CodeS: Towards Building Open-source Language Models for Text-to-SQLHaoyang Li, Jing Zhang, Hanbing Liu, Ju Fan 等SIGMOD 2024 · 被引用 124 次
- OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate SupervisionRuilin Hu, Yuyu Luo, Guoliang Li, Shuangqiao Wu 等VLDB 2026 · 被引用 4 次
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun 等VLDB 2024 · 被引用 609 次
- SDE-SQL: Enhancing Text-to-SQL Generation in Large Language Models via Self-Driven Exploration with SQL ProbesWenxuan Xie, Yaxun Dai, Wenhao JiangACL 2026 · 被引用 4 次
- Addressing Semantic Blind Spots in Text-to-SQL via Component Pre-generation and AST Matching RewardsXingyu Ma, Xin Tian, Lingxiang Wu, Xuepeng Wang 等ICML 2026
