Reward-SQL: Boosting Text-to-SQL via Stepwise Execution-Aware Reasoning and Process-Supervised Rewards
Yuxin Zhang, Meihao Fan, Ju Fan, Mingyang Yi, Yuyu Luo, Guoliang Li, Bin Wu, Wenchao Zhou
摘要
YUYU LUO, HKUST (GZ), China GUOLIANG LI, Tsinghua University, China BIN WU, Alibaba Cloud Computing, China WENCHAO ZHOU, Alibaba Cloud Computing, China Recent advances in large language models (LLMs) trained with reinforcement learning (RL) have improved Text-to-SQL performance. However, RL-based approaches still struggle with complex queries due to two key limitations: insufficient stepwise execution-aware reasoning grounded in database feedback, and the lack of process-level rewards for guiding reasoning optimization. To address these issues, we propose CoCTE, a divideand-conquer and execution-aware reasoning framework that progressively composes SQL queries through intermediate view validation and structured Common Table Expressions (CTEs), improving both accuracy and interpretability. To realize a CoCTE reasoning process, we develop Reward-SQL, a unified approach with three stages: (1) model initialization, which equips LLMs with structured CoCTE reasoning capabilities; (2) process reward design, which delivers fine-grained, execution-aware supervision; and (3) process-supervised RL and inference, which integrates process rewards into training and guides the inference stage by process rewards. This paper addresses the core challenges in Reward-SQL and makes the following contributions. We introduce a process reward model (PRM) that combines execution-aware trajectory scoring with entropy-based step weighting, providing dense and interpretable supervision across reasoning steps. We integrate PRM into both RL training and inference stages, stabilizing optimization and improving trajectory exploration with process-level signals. Experiments show that Reward-SQL significantly outperforms baselines with comparable model sizes, and exhibits strong cross-domain generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DeepPrep: An LLM-Powered Agentic System for Autonomous Data PreparationMeihao Fan, Ju Fan, Yuxin Zhang, Shaolei Zhang 等VLDB 2026 · 被引用 4 次
- TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database QueriesChao Deng, Ju Fan, Yuyu Luo, Qinliang Xue 等VLDB 2026 · 被引用 1 次
- Addressing Semantic Blind Spots in Text-to-SQL via Component Pre-generation and AST Matching RewardsXingyu Ma, Xin Tian, Lingxiang Wu, Xuepeng Wang 等ICML 2026
它引用的顶会 Paper26
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun 等VLDB 2024 · 被引用 609 次
相关 Paper
- ReEx-SQL: Reasoning with Execution-Aware Reinforcement Learning for Text-to-SQLYaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang 等ACL 2026 · 被引用 8 次
- MARS-SQL: A Multi-Agent Reinforcement Learning Framework For Text-To-SQLHaolin Yang, Jipeng Zhang, Zhitao He, Alexander Zhou 等ICML 2026 · 被引用 12 次
- Boosting Small Language Models for Text-to-SQL with Fine-Grained Execution Feedback and Cost-Efficient RewardsThanh Dat Hoang, Thanh Trung Huynh, Matthias Weidlich, Thanh Tam Nguyen 等ICDE 2026 · 被引用 3 次
- CogSQL: A Cognitive Framework for Enhancing Large Language Models in Text-to-SQL TranslationHongwei Yuan, Xiu Tang, Ke Chen, Lidan Shou 等AAAI 2025 · 被引用 12 次
- STaR-SQL: Self-Taught Reasoner for Text-to-SQLMingqian He, Yongliang Shen, Wenqi Zhang, Qiuying Peng 等ACL 2025
