BOOSTAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
Yuanhao Li, Hongbo Wang, Xiaotang Shang, Xunzhu Tang, Yiming Cao, Xuhong Chen
Abstract
Reinforcement learning for program repair is hindered by sparse execution feedback and coarse sequence-level rewards that obscure which edits actually fix bugs. We present BOOSTAPR, a three-stage framework addressing these challenges: (1) supervised fine-tuning on executionverified demonstrations with reasoning traces, ( 2 ) training dual reward models-a sequence-level assessor and a line-level credit allocator-from execution outcomes, and (3) PPO optimization where the line-level model redistributes rewards to critical edit regions. This line-level credit assignment operates at an intermediate granularity naturally suited to code changes. Trained on SWE-Gym and evaluated on four benchmarks, BOOST-APR achieves 40.7% on SWE-bench Verified (+22.9pp over base model), 24.8% on Defects4J (Python→Java transfer), 84.5% on HumanEval-Java, and 95.0% on QuixBugs, achieving competitive results among open-source models with strong cross-language generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f76be9b2-c087-4844-b396-cb33cdf5beb6Cited by top-tier papers2
- Focus-Then-Contact: Speeding Up Robotic Contact-Rich Task Learning with Affordance-Guided Real-World Residual Reinforcement LearningGuanren Qiao, Ruixiang Ouyang, Sheng Xu, Ruixing Jin et al.ICML 2026
- Deterministic Component Mining for Multi-Framework UI2Code GenerationZixiong Yang, Linxiao Li, Jiaye Lin, Binrui Wu et al.ICML 2026
Builds on19
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- SelfAPR: Self-supervised Program Repair with Test Execution DiagnosticsHe Ye, Matias Martinez, Xiapu Luo, Tao Zhang et al.ASE 2022 · 75 citations
- KNOD: Domain Knowledge Distilled Tree Decoder for Automated Program RepairNan Jiang, Thibaud Lutellier, Yiling Lou, Lin Tan et al.ICSE 2023 · 55 citations
- ThinkRepair: Self-Directed Automated Program RepairXin Yin, Chao Ni, Shaohua Wang, Zhenhao Li et al.ISSTA 2024 · 37 citations
- Debugging Engine Enhanced by Prior Knowledge: Can We Teach LLM How to Debug?Kunyi Li, Sai Wu, Xiu Tang, Chang Yao et al.FSE 2026 · 1 citation
- StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement LearningHao Wang, Lei Sha, Jie ZhangICML 2026
