An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
Wei Sun, Qianlong Du, Fuwei Cui, Jiajun Zhang
Abstract
Enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) is of great scientific and practical significance. Researchers typically employ process-supervised reward models (PRMs) to guide the reasoning process, effectively improving the models' reasoning abilities. However, existing methods for constructing process supervision training data, such as manual annotation and perstep Monte Carlo estimation, are often costly or suffer from poor quality. To address these challenges, this paper introduces a framework called EpicPRM (Efficient, Precise, Cheap), which annotates each intermediate reasoning step based on its quantified contribution and uses an adaptive binary search algorithm to enhance both annotation precision and efficiency. Using this approach, we efficiently construct a high-quality process supervision training dataset named Epic50k, consisting of 50k annotated intermediate steps. Compared to other publicly available datasets, the PRM trained on Epic50k demonstrates significantly superior performance. Getting Epic50k at https://github.com/xiaolizh1/EpicPRM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6548e97e-5f34-4f9c-b4b1-2af0bd56159fCited by top-tier papers6
- KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical ReasoningWei Sun, Wen Yang, Pu Jian, Qianlong Du et al.NeurIPS 2025 · 22 citations
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning AbilitiesJiayi Kuang, Haojing Huang, Yinghui Li, Xinnian Liang et al.NeurIPS 2025 · 11 citations
- ThoughtFold: Folding Reasoning Chains via Introspective Preference LearningZiyan Liu, Xueda Shen, Yuzhe Gu, Songyang Gao et al.ICML 2026 · 3 citations
- Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical ReasoningJiuzhou Han, Wray L. Buntine, Ehsan ShareghiAAAI 2026 · 3 citations
- Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level RewardsRaffaele Pisano, Roberto NavigliACL 2026 · 2 citations
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
Related papers
- R-PRM: Reasoning-Driven Process Reward ModelingShuaijie She, Junxiao Liu, Yifeng Liu, Jiajun Chen et al.EMNLP 2025
- Adversarial Training for Process Reward ModelsGurusha Juneja, Deepak Nathani, William WangICML 2026 · 2 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- OR-PRM: A Process Reward Model for Algorithmic Problem in Operations ResearchYilin Wang, Heng Zhou, Dongxing Mao, Linjie Li et al.ICLR 2026
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward LearningYuyang Ding, Xinyu Shi, Juntao Li, Xiaobo Liang et al.NeurIPS 2025 · 10 citations
