From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
Zhengyu Chen, Yudong Wang, Teng Xiao, Ruochen Zhou, Xuesheng Yang, Wei Wang, Zhifang Sui, Jingang Wang
Abstract
Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, including training methodologies, scalability, and generalization capabilities. We investigate the interplay between pre-training and reward model training FLOPs to assess their influence on PRM efficiency and accuracy in complex reasoning tasks. Our analysis reveals a pattern of diminishing returns in performance with increasing PRM scale, highlighting the importance of balancing model size and computational cost. Furthermore, the diversity of training datasets significantly impacts PRM performance, emphasizing the importance of diverse data to enhance both accuracy and efficiency. We further examine test-time scaling strategies, identifying Monte Carlo Tree Search as the most effective method when computational resources are abundant, while Best-of-N Sampling serves as a practical alternative under resource-limited conditions. Notably, our findings indicate that PRMs trained on mathematical datasets exhibit performance comparable to those tailored for code generation, suggesting robust cross-domain generalization. Employing a gradient-based metric, we observe that PRMs exhibit a preference for selecting responses with similar underlying patterns, further informing their optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70fe6da0-d98b-4bb4-a86b-2573d9b1d560Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- BA-GNN: On Learning Bias-Aware Graph Neural NetworkZhengyu Chen, Teng Xiao, Kun KuangICDE 2022 · 28 citations
- AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power LawsOren Neumann, Claudius GrosNeurIPS 2025 · 11 citations
- Inference Scaling for Long-Context Retrieval Augmented GenerationZhenrui Yue, Honglei Zhuang, Aijun Bai, Kai Hui et al.ICLR 2025
Related papers
- Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level RewardsRaffaele Pisano, Roberto NavigliACL 2026 · 2 citations
- VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning DataThomas Zeng, Shuibai Zhang, Shutong Wu, Christian Classen et al.ICML 2025
- Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for ReasoningCharlie Victor Snell, Jaehoon Lee, Kelvin Xu, Aviral KumarICLR 2025
- A Comprehensive Survey of Process Reward Models: Data Generation, Model Construction, and UsageCongmin Zheng, Jiachen Zhu, Zhuoying Ou, Yuxiang Chen et al.ACL 2026
- Rewarding Graph Reasoning Process makes LLMs more Generalized ReasonersMiao Peng, Nuo Chen, Zongrui Suo, Jia LiKDD 2025 · 1 citation
