Multi-Turn Code Generation Through Single-Step Rewards
Arnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen, Alexander M. Rush, Wenting Zhao, Sanjiban Choudhury
摘要
We address the problem of code generation from multi-turn execution feedback. Existing methods either generate code without feedback or use complex, hierarchical reinforcement learning to optimize multi-turn rewards. We propose a simple yet scalable approach, µCODE, that solves multi-turn code generation using only single-step rewards. Our key insight is that code generation is a one-step recoverable MDP, where the correct code can be recovered from any intermediate code state in a single turn. µCODE iteratively trains both a generator to provide code solutions conditioned on multi-turn execution feedback and a verifier to score the newly generated code. Experimental evaluations show that our approach achieves significant improvements over the stateof-the-art baselines. We provide analysis of the design choices of the reward models and policy, and show the efficacy of µCODE at utilizing the execution feedback. Our code is available here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic TasksShuo He, Lang Feng, Qi Wei, Xin Cheng 等ICLR 2026 · 被引用 36 次
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to SearchArnav Kumar Jain, Vibhakar Mohta, Subin Kim, Atiksh Bhardwaj 等NeurIPS 2025 · 被引用 27 次
- V1: Unifying Generation and Self-Verification for Parallel ReasonersHarman Singh, Xiuyu Li, Kusha Sareen, Monishwaran Maheswaran 等ICML 2026 · 被引用 8 次
- RedCoder: Automated Multi-Turn Red Teaming for Code LLMsWenjie Jacky Mo, Qin Liu, Xiaofei Wen, Dongwon Jung 等ACL 2026 · 被引用 6 次
- MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-CorrectionZitian Tang, Xu Zhang, Jianbo Yuan, Yang Zou 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
相关 Paper
- ReVeal: Self-Evolving Code Agents via Reliable Self-VerificationYiyang Jin, Kunzhao Xu, Hang Li, Xueting Han 等ICLR 2026 · 被引用 13 次
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella 等ICML 2025
- CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process SupervisionYifei Lu, Fanghua Ye, Jian Li, Qiang Gao 等ACL 2025 · 被引用 8 次
- ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution ReasoningLingxiao Tang, He Ye, Zhaoyang Chu, Muyang Ye 等ACL 2026 · 被引用 2 次
- StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement LearningHao Wang, Lei Sha, Jie ZhangICML 2026
