Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning
Yinjie Wang, Ling Yang, Ye Tian, Ke Shen, Mengdi Wang
Abstract
Mathematical reasoning in large language models has been successfully incentivized through reinforcement learning with verifiable rewards, leading to improved one-shot precision. In this work, we turn our focus to the coding domain. Beyond one-shot precision, we highlight unit test generation as another key factor for enhancing coding ability, since accurate unit tests are essential for enabling self-checking and self-correction during inference. Traditional approaches for finetuning LLMs on unit test generation rely heavily on ground-truth code solutions in the training data. We propose CURE, a novel reinforcement learning framework with a dedicated reward design that co-evolves coding and unit test generation capabilities based on their interaction outcomes-without any ground-truth code as supervision. This approach enables flexible and scalable training and allows the unit tester to learn directly from the coder's mistakes. Through extensive evaluations, we demonstrate that our CURE models, derived from base models of varying sizes, excel in both code generation and unit test generation. They naturally extend to downstream tasks such as test-time scaling-achieving a 6.2% improvement over the base model-and agentic unit test generation, with a 25.1% improvement. Our CURE-4B model consistently outperforms Qwen3-4B while achieving 64.8% inference efficiency in unit test generation. Notably, we also find that the CURE model can serve as an effective reward model for reinforcement learning on base models, even in the absence of any labeled supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee474a38-ebf8-4d6c-8270-aa5cb1ec7b6bCited by top-tier papers19
- R-Zero: Self-Evolving Reasoning LLM from Zero DataChengsong Huang, Wenhao Yu, Xiaoyang Wang, Hongming Zhang et al.ICLR 2026 · 220 citations
- Revolutionizing Reinforcement Learning Framework for Diffusion Large Language ModelsYinjie Wang, Ling Yang, Bowen Li, Ye Tian et al.ICLR 2026 · 79 citations
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMsJiaru Zou, Ling Yang, Jingwen Gu, Jiahao Qiu et al.NeurIPS 2025 · 51 citations
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language ModelsMickel Liu, Liwei Jiang, Yancheng Liang, Simon Du et al.ICML 2026 · 34 citations
- Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMsYujie Zhao, Lanxiang Hu, Yang Wang, Minmin Hou et al.ICLR 2026 · 26 citations
Builds on18
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
Related papers
- CURE: Critique-Driven Unified Reinforcement Learning for Test-Time Self-ImprovementGuirong Chen, Shuqi Ye, Wenkai Yang, Shiqi Shen et al.ACL 2026
- Learning to Generate Unit Test via Adversarial Reinforcement LearningDongjun Lee, Changho Hwang, Kimin LeeICLR 2026 · 14 citations
- ATGen: Adversarial Reinforcement Learning for Test Case GenerationQingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li et al.ICLR 2026 · 4 citations
- Dynamic Scaling of Unit Tests for Code Reward ModelingZeyao Ma, Xiaokang Zhang, Jing Zhang, Jifan Yu et al.ACL 2025 · 21 citations
- Anchoring Self-Play for Code RepairCaroline Choi, Zeyneb Kaya, Shirley Wu, Tengyu Ma et al.ICML 2026 · 1 citation
