Process Reward Model with Q-value Rankings
Wendi Li, Yixuan Li
摘要
Process Reward Modeling (PRM) is critical for complex reasoning and decisionmaking tasks where the accuracy of intermediate steps significantly influences the overall outcome. Existing PRM approaches, primarily framed as classification problems, employ cross-entropy loss to independently evaluate each step's correctness. This method can lead to suboptimal reward distribution and does not adequately address the interdependencies among steps. To address these limitations, we introduce the Process Q-value Model (PQM), a novel framework that redefines PRM in the context of a Markov Decision Process. PQM optimizes Qvalue rankings based on a novel comparative loss function, enhancing the model's ability to capture the intricate dynamics among sequential decisions. This approach provides a more granular and theoretically grounded methodology for process rewards. Our extensive empirical evaluations across various sampling policies, language model backbones, and multi-step reasoning benchmarks show that PQM outperforms classification-based PRMs. The effectiveness of the comparative loss function is highlighted in our comprehensive ablation studies, confirming PQM's practical efficacy and theoretical advantage. Our codes can be found at https://github.com/WindyLee0822/Process Q Model .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen 等NeurIPS 2025 · 被引用 177 次
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningJian Zhao, Runze Liu, Kaiyan Zhang, Zhimu Zhou 等AAAI 2026 · 被引用 68 次
- Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced ReasoningShaokun Zhang, Yi Dong, Jieyu Zhang, Jan Kautz 等ICLR 2026 · 被引用 61 次
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMsJiaru Zou, Ling Yang, Jingwen Gu, Jiahao Qiu 等NeurIPS 2025 · 被引用 51 次
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
相关 Paper
- Linking Process to Outcome: Conditional Reward Modeling for LLM ReasoningZheng Zhang, Ziwei Shan, Kaitao Song, Yexin Li 等ICLR 2026 · 被引用 16 次
- Discriminative Policy Optimization for Token-Level Reward ModelsHongzhan Chen, Tao Yang, Shiping Gao, Ruijun Chen 等ICML 2025
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained RewardsHonghao Chen, Xingzhou Lou, Xiaokun Feng, Kaiqi Huang 等NeurIPS 2025 · 被引用 7 次
- PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward ModelsMingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou 等ACL 2025 · 被引用 85 次
- The Bidirectional Process Reward ModelLingyin Zhang, Jun Gao, Xiaoxue Ren, Ziqiang CaoACL 2026 · 被引用 2 次
