Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
Miao Peng, Nuo Chen, Zongrui Suo, Jia Li
Abstract
Despite significant advancements in Large Language Models (LLMs), developing advanced reasoning capabilities in LLMs remains a key challenge. Process Reward Models (PRMs) have demonstrated exceptional promise in enhancing reasoning by providing step-wise feedback, particularly in the context of mathematical reasoning. However, their application to broader reasoning domains remains understudied, largely due to the high costs associated with manually creating step-level supervision. In this work, we explore the potential of PRMs in graph reasoning problems - a domain that demands sophisticated multi-step reasoning and offers opportunities for automated step-level data generation using established graph algorithms. We introduce GraphSilo, the largest dataset for graph reasoning problems with fine-grained step-wise label, built using automated Task-oriented Trajectories and Monte Carlo Tree Search (MCTS) to generate detailed reasoning steps with step-wise labels. Building upon this dataset, we train GraphPRM, the first PRM designed for graph reasoning problems, and evaluate its effectiveness in two key settings: inference-time scaling and reinforcement learning via Direct Preference Optimization (DPO). Experimental results show that GraphPRM significantly improves LLM performance across 13 graph reasoning tasks, delivering a 9% gain for Qwen2.5-7B and demonstrating transferability to new graph reasoning datasets and new reasoning domains like mathematical problem-solving. Notably, GraphPRM enhances LLM performance on GSM8K and MATH500, underscoring the cross-domain applicability of graph-based reasoning rewards. Our findings highlight the potential of PRMs in advancing reasoning across diverse domains, paving the way for more versatile and effective LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71c7ab17-6db3-4fbb-a7c7-f348bf3b40ecCited by top-tier papers7
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?Junchi Yu, Yujie Liu, Jindong Gu, Philip H. S. Torr et al.NeurIPS 2025 · 8 citations
- Chain of Execution Supervision Promotes General Reasoning in Large Language ModelsNuo Chen, Zehua Li, Keqin Bao, Junyang Lin et al.NeurIPS 2025 · 6 citations
- Exposing Weaknesses of Large Reasoning Models through Graph Algorithm ProblemsQifan Zhang, Jianhao Ruan, Aochuan Chen, Kang Zeng et al.ICLR 2026 · 4 citations
- From Sequence to Structure: Uncovering Substructure Reasoning in TransformersXinnan Dai, Kai Yang, Jay Revolinsky, Kai Guo et al.NeurIPS 2025 · 3 citations
- RouteGoT: Node-Adaptive Routing for Cost-Efficient Graph of Thoughts ReasoningYuhang Liu, Ruijie Wang, Yunlong Chu, Bing Hao et al.KDD 2026 · 1 citation
Builds on26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level RewardsRaffaele Pisano, Roberto NavigliACL 2026 · 2 citations
- From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time ScalingZhengyu Chen, Yudong Wang, Teng Xiao, Ruochen Zhou et al.AAAI 2026 · 2 citations
- <tt>G1</tt>: Teaching LLMs to Reason on Graphs with Reinforcement LearningXiaojun Guo, Ang Li, Yifei Wang, Stefanie Jegelka et al.NeurIPS 2025 · 16 citations
- VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning DataThomas Zeng, Shuibai Zhang, Shutong Wu, Christian Classen et al.ICML 2025
- Discriminative Policy Optimization for Token-Level Reward ModelsHongzhan Chen, Tao Yang, Shiping Gao, Ruijun Chen et al.ICML 2025
