StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
Shihan Dou, Yan Liu, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou, Tao Ji
Abstract
The advancement of large language models (LLMs) has significantly propelled the field of code generation. Previous work integrated reinforcement learning (RL) with compiler feedback for exploring the output space of LLMs to enhance code generation quality. However, the lengthy code generated by LLMs in response to complex human requirements makes RL exploration a challenge. Also, since the unit tests may not cover the complicated code, optimizing LLMs by using these unexecuted code snippets is ineffective. To tackle these challenges, we introduce StepCoder, a novel RL framework for code generation, consisting of two main components: CCCS addresses the exploration challenge by breaking the long sequences code generation task into a Curriculum of Code Completion Subtasks, while FGO only optimizes the model by masking the unexecuted code segments to provide Fine-Grained Optimization. In addition, we furthermore construct the APPS+ dataset for RL training, which is manually verified to ensure the correctness of unit tests. Experimental results show that our method improves the ability to explore the output space and outperforms state-of-the-art approaches in corresponding benchmarks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef90d6d1-5b8f-4acc-bb03-9bfeb1cb6496Cited by top-tier papers10
- Agentic Reinforcement Learning with Implicit Step RewardsXiaoqian Liu, Ke Wang, Yuchuan Wu, Fei Huang et al.ICLR 2026 · 46 citations
- ReCode: Reinforcing Code Generation with Reasoning-Process RewardsLishui Fan, Yu Zhang, Mouxiang Chen, Zhongxin LiuACL 2026 · 19 citations
- CODERL+: Improving Code Generation via Reinforcement with Execution Semantics AlignmentXue Jiang, Yihong Dong, Mengyang Liu, Hongyi Deng et al.ACL 2026 · 18 citations
- Dynamic and Generalizable Process Reward ModelingZhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Qinyuan Cheng et al.ACL 2025 · 13 citations
- Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQLHanbing Liu, Haoyang Li, Xiaokang Zhang, Ruotong Chen et al.ACL 2025 · 10 citations
Builds on2
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese et al.NeurIPS 2022 · 571 citations
Related papers
- Learning to Generate Unit Test via Adversarial Reinforcement LearningDongjun Lee, Changho Hwang, Kimin LeeICLR 2026 · 14 citations
- SeDev: Structured Semantic Exploration for LLM-Driven Code GenerationRonghui Yang, Jie Liu, Jiajie Zeng, Jiexin Wang et al.ACL 2026
- Alignment with Fill-In-the-Middle for Enhancing Code GenerationHouxing Ren, Zimu Lu, Weikang Shi, Haotian Hou et al.EMNLP 2025
- UnitCoder: Scalable Code Synthesis from Pre-training CorporaYichuan Ma, Yunfan Shao, Peiji Li, Demin Song et al.EMNLP 2025 · 2 citations
- JumpCoder: Go Beyond Autoregressive Coder via Online ModificationMouxiang Chen, Hao Tian, Zhongxin Liu, Xiaoxue Ren et al.ACL 2024 · 4 citations
