Dynamic Scaling of Unit Tests for Code Reward Modeling
Zeyao Ma, Xiaokang Zhang, Jing Zhang, Jifan Yu, Sijia Luo, Jie Tang
Abstract
Current large language models (LLMs) often struggle to produce accurate responses on the first attempt for complex reasoning tasks like code generation. Prior research tackles this challenge by generating multiple candidate solutions and validating them with LLM-generated unit tests. The execution results of unit tests serve as reward signals to identify correct solutions. As LLMs always confidently make mistakes, these unit tests are not reliable, thereby diminishing the quality of reward signals. Motivated by the observation that scaling the number of solutions improves LLM performance, we explore the impact of scaling unit tests to enhance reward signal quality. Our pioneer experiment reveals a positive correlation between the number of unit tests and reward signal quality, with greater benefits observed in more challenging problems. Based on these insights, we propose CodeRM-8B, a lightweight yet effective unit test generator that enables efficient and high-quality unit test scaling. Additionally, we implement a dynamic scaling mechanism that adapts the number of unit tests based on problem difficulty, further improving efficiency. Experimental results show that our approach significantly improves performance across various models on three benchmarks (e.g., with gains of 18.43% for Llama3-8B and 3.42% for GPT-4o-mini on HumanEval Plus).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8068e2a-7fa6-47ff-9386-2c7747fd949bCited by top-tier papers9
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement LearningSiyan Zhao, Devaansh Gupta, Qinqing Zheng, Aditya GroverNeurIPS 2025 · 191 citations
- Co-Evolving LLM Coder and Unit Tester via Reinforcement LearningYinjie Wang, Ling Yang, Ye Tian, Ke Shen et al.NeurIPS 2025 · 56 citations
- Rethinking Verification for LLM Code Generation: From Generation to TestingZihan Ma, Taolin Zhang, Maosong Cao, Junnan Liu et al.NeurIPS 2025 · 19 citations
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement LearningRan Xu, Jingjing Chen, Jiayu Ye, Yu Wu et al.ICLR 2026 · 17 citations
- Learning to Generate Unit Test via Adversarial Reinforcement LearningDongjun Lee, Changho Hwang, Kimin LeeICLR 2026 · 14 citations
Builds on14
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Generative Verifiers: Reward Modeling as Next-Token PredictionLunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi et al.ICLR 2025
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
- CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksHongchao Jiang, Yiming Chen, Yushi Cao, Hung-Yi Lee et al.ACL 2026 · 33 citations
- UnitCoder: Scalable Code Synthesis from Pre-training CorporaYichuan Ma, Yunfan Shao, Peiji Li, Demin Song et al.EMNLP 2025 · 2 citations
- Oracle-Guided Program Selection from Large Language ModelsZhiyu Fan, Haifeng Ruan, Sergey Mechtaev, Abhik RoychoudhuryISSTA 2024 · 4 citations
