B-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis
Zishun Yu, Yunzhe Tao, Liyu Chen, Tao Sun, Hongxia Yang
Abstract
Program synthesis aims to create accurate, executable programs from problem specifications, specifically from natural language descriptions in our context. Recent studies have leveraged the power of reinforcement learning (RL) in conjunction with large language models (LLMs), significantly enhancing code generation capabilities. The application of RL focuses on directly optimizing for functional correctness, offering an advantage over conventional supervised methods. Despite policy-based RL methods dominating the literature on RL for program synthesis, the nature of program synthesis tasks hints at a natural alignment with value-based methods. This stems from the rich collection of off-policy programs, including those developed by human programmers and also historical samples, coupled with the straightforward verification of generated programs through automated unit testing, meaning rewards are easy to obtain. Diverging from the dominant use of policy-based algorithms, our work explores the feasibility of value-based approaches, leading to the development of our -Coder (pronounced Bellman coder). Yet, training value-based methods presents challenges due to the enormous search space inherent to program synthesis. To this end, we introduce an initialization protocol for RL agents utilizing pre-trained LMs and a conservative Bellman operator to reduce training complexities. Moreover, we demonstrate how to leverage the learned value functions as a dual strategy to post-process generated programs. Our empirical evaluations demonstrated -Coder's capability in achieving state-of-the-art performance when compared to policy-based methods. Remarkably, this achievement is reached with minimal reward engineering effort, highlighting the effectiveness of value-based RL, independent of reward designs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 367e7bf2-22ee-494a-b4a2-ed6f227b9a06Cited by top-tier papers7
- Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning ModelsSoumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu et al.NeurIPS 2025 · 43 citations
- Training Language Models to Generate Quality Code with Program Analysis FeedbackFeng Yao, Zilong Wang, Liyuan Liu, Junxia Cui et al.NeurIPS 2025 · 11 citations
- NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding ModificationWenshuo Zhang, Leixian Shen, Shuchang Xu, Jindu Wang et al.UIST 2025 · 9 citations
- AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph ConstructionHong Ting Tsang, Jiaxin Bai, Haoyu Huang, Qiao Xiao et al.ACL 2026 · 4 citations
- Reinforcement Learning-Guided Data Selection Via Redundancy AssessmentSuorong Yang, Peijia Li, Furao Shen, Jian ZhaoICCV 2025 · 1 citation
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
Related papers
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella et al.ICML 2025
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese et al.NeurIPS 2022 · 571 citations
- ACECODER: Acing Coder RL via Automated Test-Case SynthesisHuaye Zeng, Dongfu Jiang, Haozhe Wang, Ping Nie et al.ACL 2025 · 72 citations
- Code as Reward: Empowering Reinforcement Learning with VLMsDavid Venuto, Mohammad Sami Nur Islam, Martin Klissarov, Doina Precup et al.ICML 2024 · 29 citations
- Process-Supervised Reinforcement Learning for Code GenerationYufan Ye, Ting Zhang, Wenbin Jiang, Hua HuangEMNLP 2025 · 1 citation
