Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
Raffaele Pisano, Roberto Navigli
Abstract
Process Reward Models (PRMs) have emerged as a powerful tool for providing step-level feedback when evaluating the reasoning of Large Language Models (LLMs), which frequently produce chains of thought (CoTs) containing errors even when the final answer is correct. However, existing PRM datasets remain expensive to construct, prone to annotation errors, and predominantly limited to the mathematical domain. This work introduces a novel and scalable approach to PRM dataset generation based on planning logical problems expressed in the Planning Domain Definition Language (PDDL). Using this method, we generate a corpus of approximately one million reasoning steps across various PDDL domains and use it to train PRMs. Experimental results show that augmenting widely-used PRM training datasets with PDDL-derived data yields substantial improvements in both mathematical and non-mathematical reasoning, as demonstrated across multiple benchmarks. These findings indicate that planning problems constitute a scalable and effective resource for generating robust, precise, and fine-grained training data for PRMs, going beyond the classical mathematical sources that dominate this field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- On the Planning Abilities of Large Language Models - A Critical InvestigationKarthik Valmeekam, Matthew Marquez, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 509 citations
- Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task PlanningLin Guan, Karthik Valmeekam, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 347 citations
Related papers
- ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time ScalingHaotian Zhang, Liu Liu, Baosheng Yu, Jiayan Qiu et al.ICLR 2026
- From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time ScalingZhengyu Chen, Yudong Wang, Teng Xiao, Ruochen Zhou et al.AAAI 2026 · 2 citations
- R-PRM: Reasoning-Driven Process Reward ModelingShuaijie She, Junxiao Liu, Yifeng Liu, Jiajun Chen et al.EMNLP 2025
- VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning DataThomas Zeng, Shuibai Zhang, Shutong Wu, Christian Classen et al.ICML 2025
- OR-PRM: A Process Reward Model for Algorithmic Problem in Operations ResearchYilin Wang, Heng Zhou, Dongxing Mao, Linjie Li et al.ICLR 2026
