Multi-Turn Code Generation Through Single-Step Rewards
Arnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen, Alexander M. Rush, Wenting Zhao, Sanjiban Choudhury
Abstract
We address the problem of code generation from multi-turn execution feedback. Existing methods either generate code without feedback or use complex, hierarchical reinforcement learning to optimize multi-turn rewards. We propose a simple yet scalable approach, µCODE, that solves multi-turn code generation using only single-step rewards. Our key insight is that code generation is a one-step recoverable MDP, where the correct code can be recovered from any intermediate code state in a single turn. µCODE iteratively trains both a generator to provide code solutions conditioned on multi-turn execution feedback and a verifier to score the newly generated code. Experimental evaluations show that our approach achieves significant improvements over the stateof-the-art baselines. We provide analysis of the design choices of the reward models and policy, and show the efficacy of µCODE at utilizing the execution feedback. Our code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13bbe22a-d716-4899-9b16-86e7bb656b08Cited by top-tier papers9
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic TasksShuo He, Lang Feng, Qi Wei, Xin Cheng et al.ICLR 2026 · 36 citations
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to SearchArnav Kumar Jain, Vibhakar Mohta, Subin Kim, Atiksh Bhardwaj et al.NeurIPS 2025 · 27 citations
- V1: Unifying Generation and Self-Verification for Parallel ReasonersHarman Singh, Xiuyu Li, Kusha Sareen, Monishwaran Maheswaran et al.ICML 2026 · 8 citations
- RedCoder: Automated Multi-Turn Red Teaming for Code LLMsWenjie Jacky Mo, Qin Liu, Xiaofei Wen, Dongwon Jung et al.ACL 2026 · 6 citations
- MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-CorrectionZitian Tang, Xu Zhang, Jianbo Yuan, Yang Zou et al.CVPR 2026 · 4 citations
Builds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
Related papers
- ReVeal: Self-Evolving Code Agents via Reliable Self-VerificationYiyang Jin, Kunzhao Xu, Hang Li, Xueting Han et al.ICLR 2026 · 13 citations
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella et al.ICML 2025
- CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process SupervisionYifei Lu, Fanghua Ye, Jian Li, Qiang Gao et al.ACL 2025 · 8 citations
- ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution ReasoningLingxiao Tang, He Ye, Zhaoyang Chu, Muyang Ye et al.ACL 2026 · 2 citations
- StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement LearningHao Wang, Lei Sha, Jie ZhangICML 2026
