R-PRM: Reasoning-Driven Process Reward Modeling
Shuaijie She, Junxiao Liu, Yifeng Liu, Jiajun Chen, Xin Huang, Shujian Huang
摘要
Process Reward Models (PRMs) have emerged as a promising solution to address the reasoning mistakes of large language models (LLMs). However, existing PRMs typically output evaluation scores directly, limiting both learning efficiency and evaluation accuracy. This limitation is further compounded by the scarcity of annotated data. To address these issues, we propose Reasoning-Driven Process Reward Modeling (R-PRM), which activates inherent reasoning to enhance process-level evaluation. First, we leverage stronger LLMs to generate seed data from limited annotations, effectively activating reasoning capabilities and enabling comprehensive step-by-step evaluation. Second, we explore self-improvement of our PRM through preference optimization, without requiring additional annotated data. Third, we introduce inference time scaling to fully harness our model's reasoning potential. Extensive experiments demonstrate R-PRM's effectiveness: on ProcessBench and PRMBench, it surpasses strong baselines by 13.9 and 8.5 F1 scores. When applied to guide mathematical reasoning, R-PRM achieves consistent accuracy improvements of over 8.6 points across six challenging datasets. Further analysis reveals that R-PRM exhibits more comprehensive evaluation and robust generalization, indicating its broader potential.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Optimas: Optimizing Compound AI Systems with Globally Aligned Local RewardsShirley Wu, Parth Sarthi, Shiyu Zhao, Aaron Lee 等ICLR 2026 · 被引用 25 次
- Linking Process to Outcome: Conditional Reward Modeling for LLM ReasoningZheng Zhang, Ziwei Shan, Kaitao Song, Yexin Li 等ICLR 2026 · 被引用 16 次
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMsJingyao Wang, Wenwen Qiang, Zeen Song, Changwen Zheng 等NeurIPS 2025 · 被引用 13 次
- SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward ModellingMd Imbesat Hassan Rizvi, Xiaodan Zhu, Iryna GurevychAAAI 2026 · 被引用 2 次
- From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level GroundingYuyuan Liu, Yiping Ji, Anjie Le, Jiayuan Zhu 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin 等ACL 2025 · 被引用 209 次
相关 Paper
- Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical ReasoningJiuzhou Han, Wray L. Buntine, Ehsan ShareghiAAAI 2026 · 被引用 3 次
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward LearningYuyang Ding, Xinyu Shi, Juntao Li, Xiaobo Liang 等NeurIPS 2025 · 被引用 10 次
- Adversarial Training for Process Reward ModelsGurusha Juneja, Deepak Nathani, William WangICML 2026 · 被引用 2 次
- Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMsZhangyin Feng, Qianglong Chen, Ning Lu, Yongqian Li 等NeurIPS 2025 · 被引用 16 次
- Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level RewardsRaffaele Pisano, Roberto NavigliACL 2026 · 被引用 2 次
