Dynamic and Generalizable Process Reward Modeling
Zhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Qinyuan Cheng, Xipeng Qiu, Xuanjing Huang
Abstract
Process Reward Models (PRMs) are crucial for guiding Large Language Models (LLMs) in complex scenarios by providing dense reward signals. However, existing PRMs primarily rely on heuristic approaches, which struggle with cross-domain generalization. While LLM-as-judge has been proposed to provide generalized rewards, current research has focused mainly on feedback results, overlooking the meaningful guidance embedded within the text. Additionally, static and coarse-grained evaluation criteria struggle to adapt to complex process supervision. To tackle these challenges, we propose Dynamic and Generalizable Process Reward Modeling (DG-PRM), which features a reward tree to capture and store fine-grained, multi-dimensional reward criteria. DG-PRM dynamically selects reward signals for step-wise reward scoring. To handle multifaceted reward signals, we pioneeringly adopt Pareto dominance estimation to identify discriminative positive and negative pairs. Experimental results show that DG-PRM achieves stunning performance on prevailing benchmarks, significantly boosting model performance across tasks with dense rewards. Further analysis reveals that DG-PRM adapts well to out-of-distribution scenarios, demonstrating exceptional generalizability. "Judgements prevent us from seeing the good that lies beyond appearances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c4c15d4-376c-4df1-b38c-d5e0b6fcd596Cited by top-tier papers3
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning SegmentationZhenyu Lu, Liupeng Li, Jinpeng Wang, Yan Feng et al.ICLR 2026 · 10 citations
- From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level GroundingYuyuan Liu, Yiping Ji, Anjie Le, Jiayuan Zhu et al.CVPR 2026 · 2 citations
- SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action ModelsHyeonbeom Choi, Daechul Ahn, Youhan Lee, Taewook Kang et al.ICML 2026 · 2 citations
Builds on24
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
Related papers
- AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and ProgressZhiheng Xi, Chenyang Liao, Guanyu Li, Zhihao Zhang et al.WWW 2026 · 19 citations
- R-PRM: Reasoning-Driven Process Reward ModelingShuaijie She, Junxiao Liu, Yifeng Liu, Jiajun Chen et al.EMNLP 2025
- Linking Process to Outcome: Conditional Reward Modeling for LLM ReasoningZheng Zhang, Ziwei Shan, Kaitao Song, Yexin Li et al.ICLR 2026 · 16 citations
- Learning Ordinal Probabilistic Reward from PreferencesLongze Chen, Lu Wang, Renke Shan, Ze Gong et al.ICLR 2026 · 3 citations
- AdaJudge: Adaptive Multi-Perspective Judging for Reward ModelingYongliang Miao, Yangyang Liang, Mengnan DuACL 2026 · 1 citation
