Text2sql-Flow: a Robust Sql-Aware Data Augmentation Framework for Text-To-Sql
Qifeng Cai, Hao Liang, Chang Xu, Tao Xie, Wentao Zhang, Bin Cui
摘要
The data-centric paradigm has emerged as a pivotal direction in artificial intelligence (AI), relying on high-quality training data. This shift is especially critical in the Text-to-SQL task, where the scarcity, limited diversity, and structural simplicity of existing datasets constrain model performance. To address these challenges, we propose TEXT2SQL-FLOW, a SQLaware data augmentation framework that systematically generates large-scale, semantically valid, and structurally diverse Textto-SQL pairs from limited seed data. Our framework operates along six augmentation dimensions and integrates an end-to-end pipeline featuring auxiliary database selection, SQL executability verification, natural language (NL) question generation, NL-SQL correspondence verification, and chain-of-thought (CoT) reasoning trace generation. Leveraging this framework, we construct SQLFLOW, a high-quality dataset comprising 75,386 annotated examples. We demonstrate the utility of SQLFLOW in both finetuning and prompt-based settings: (1) For open-source large language models (LLMs), fine-tuning with SQLFLOW enhances the problem-solving capabilities. Under the same data budget, models trained on SQLFLOW achieve competitive performance gains across multiple benchmarks. (2) For closed-source LLMs, we propose a masked alignment retrieval method that leverages SQLFLOW as both a knowledge base and the training data for the retrieval model. This approach enables structure-aware example matching by modeling fine-grained alignments between NL questions and SQL queries. Experimental results show that our retrieval strategy outperforms existing example retrieval methods, highlighting the dual importance of SQLFLOW's highquality data and our novel retrieval technique. Our work establishes a scalable, data-centric foundation for advancing Textto-SQL systems and underscores the indispensable role of structured, high-fidelity data in modern AI development. Our code is available at https://github.com/TechNomad-ds/Text2SQL-Flow.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun 等VLDB 2024 · 被引用 609 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
相关 Paper
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 被引用 9 次
- SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQLJimin Lee, Ingeol Baek, Byeongjeong Kim, Hyunkyung Bae 等EMNLP 2025 · 被引用 1 次
- OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate SupervisionRuilin Hu, Yuyu Luo, Guoliang Li, Shuangqiao Wu 等VLDB 2026 · 被引用 4 次
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang 等VLDB 2025 · 被引用 90 次
- MARS-SQL: A Multi-Agent Reinforcement Learning Framework For Text-To-SQLHaolin Yang, Jipeng Zhang, Zhitao He, Alexander Zhou 等ICML 2026 · 被引用 12 次
