MARS-SQL: A Multi-Agent Reinforcement Learning Framework For Text-To-SQL
Haolin Yang, Jipeng Zhang, Zhitao He, Alexander Zhou, Yi Fung
摘要
Large Language Models (LLMs) often struggle with the precise logic and schema alignment required for complex Text-to-SQL tasks. While current methods rely heavily on static prompting, they lack the ability to dynamically adapt and self-correct through environmental interaction. To bridge this gap, we propose MARS-SQL , a trainable multi-agent framework for Text-to-SQL. Rather than introducing a new standalone SQL primitive, MARS-SQL makes an agentic workflow trainable by decomposing the problem into three specialized roles: schema grounding, query generation, and solution validation. Central to our approach is a generation agent trained via a multi-turn RL policy within a ReAct-style loop. The agent learns to iteratively reason, execute intermediate SQL actions on a live database, and refine its strategy based on execution feedback. To improve robustness, we further introduce a validation mechanism that treats solution selection as a generative modeling task, identifying the optimal interaction trajectory through next-token prediction probabilities. Empirical evaluations demonstrate the effectiveness of coupling interactive learning with trajectory ranking. MARS-SQL achieves state-of-the-art performance, recording an execution accuracy of 77.84% on the BIRD development dataset and 89.75% on the Spider test dataset, while also transferring strongly to out-of-domain benchmarks. Code is available at https://github.com/YangHaolin0526/MARS-SQL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Dancing in Chains: Strategic Persuasion in Academic Rebuttal via Theory of MindZhitao He, Zongwei Lyu, Yi R. FungICLR 2026 · 被引用 5 次
- ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic EducationZhitao He, Haolin Yang, Zeyu Qin, Yi FungICML 2026 · 被引用 4 次
- On Stable Long-Form Generation: Benchmarking and Mitigating Length VolatilityZhitao He, Haolin Yang, Rui Min, Zeyu Qin 等ICML 2026
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe 等EMNLP 2022 · 被引用 634 次
相关 Paper
- SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQLHarper Hua, Zhen Han, Zhengyuan Shen, Meng-Chieh Lee 等ACL 2026 · 被引用 3 次
- ReEx-SQL: Reasoning with Execution-Aware Reinforcement Learning for Text-to-SQLYaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang 等ACL 2026 · 被引用 8 次
- APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQLBowen Cao, Weibin Liao, Yushi Sun, Dong Fang 等KDD 2026 · 被引用 7 次
- CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQLMohammadreza Pourreza, Hailong Li, Ruoxi Sun, Yeounoh Chung 等ICLR 2025
- OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency AlignmentXiangjin Xie, Guangwei Xu, Lingyan Zhao, Ruijie GuoSIGMOD 2025 · 被引用 28 次
