R-PRM: Reasoning-Driven Process Reward Modeling
Shuaijie She, Junxiao Liu, Yifeng Liu, Jiajun Chen, Xin Huang, Shujian Huang
Abstract
Process Reward Models (PRMs) have emerged as a promising solution to address the reasoning mistakes of large language models (LLMs). However, existing PRMs typically output evaluation scores directly, limiting both learning efficiency and evaluation accuracy. This limitation is further compounded by the scarcity of annotated data. To address these issues, we propose Reasoning-Driven Process Reward Modeling (R-PRM), which activates inherent reasoning to enhance process-level evaluation. First, we leverage stronger LLMs to generate seed data from limited annotations, effectively activating reasoning capabilities and enabling comprehensive step-by-step evaluation. Second, we explore self-improvement of our PRM through preference optimization, without requiring additional annotated data. Third, we introduce inference time scaling to fully harness our model's reasoning potential. Extensive experiments demonstrate R-PRM's effectiveness: on ProcessBench and PRMBench, it surpasses strong baselines by 13.9 and 8.5 F1 scores. When applied to guide mathematical reasoning, R-PRM achieves consistent accuracy improvements of over 8.6 points across six challenging datasets. Further analysis reveals that R-PRM exhibits more comprehensive evaluation and robust generalization, indicating its broader potential.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eacfcd93-5e52-473a-a2f5-8dedb0f41742Cited by top-tier papers7
- Optimas: Optimizing Compound AI Systems with Globally Aligned Local RewardsShirley Wu, Parth Sarthi, Shiyu Zhao, Aaron Lee et al.ICLR 2026 · 25 citations
- Linking Process to Outcome: Conditional Reward Modeling for LLM ReasoningZheng Zhang, Ziwei Shan, Kaitao Song, Yexin Li et al.ICLR 2026 · 16 citations
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMsJingyao Wang, Wenwen Qiang, Zeen Song, Changwen Zheng et al.NeurIPS 2025 · 13 citations
- SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward ModellingMd Imbesat Hassan Rizvi, Xiaodan Zhu, Iryna GurevychAAAI 2026 · 2 citations
- From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level GroundingYuyuan Liu, Yiping Ji, Anjie Le, Jiayuan Zhu et al.CVPR 2026 · 2 citations
Builds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin et al.ACL 2025 · 209 citations
Related papers
- Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical ReasoningJiuzhou Han, Wray L. Buntine, Ehsan ShareghiAAAI 2026 · 3 citations
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward LearningYuyang Ding, Xinyu Shi, Juntao Li, Xiaobo Liang et al.NeurIPS 2025 · 10 citations
- Adversarial Training for Process Reward ModelsGurusha Juneja, Deepak Nathani, William WangICML 2026 · 2 citations
- Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMsZhangyin Feng, Qianglong Chen, Ning Lu, Yongqian Li et al.NeurIPS 2025 · 16 citations
- Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level RewardsRaffaele Pisano, Roberto NavigliACL 2026 · 2 citations
