AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
Yuliang Liu, Junjie Lu, Chaofeng Qu, Zhaoling Chen, Zefan Cai, Jason Klein Liu, Chonghan Liu, Yunhui Xia, Li Zhao, Jiang Bian, Chuheng Zhang, Wei Shen, Zhouhan Lin
Abstract
Current approaches for training Process Reward Models (PRMs) often involve decomposing responses into multiple reasoning steps using rulebased techniques, such as using predefined placeholder tokens or setting the reasoning step's length to a fixed size. These approaches overlook the fact that certain words don't usually indicate true decision points. To address this, we propose AdaptiveStep, a method that divides reasoning steps based on the model's confidence in predicting the next word, offering more information on decision-making at each step, improving downstream tasks like reward model training. Moreover, our method requires no manual annotation. Experiments with AdaptiveStep-trained PRMs in mathematical reasoning and code generation show that the outcome PRM achieves state-of-the-art Best-of-N performance, surpassing greedy search strategy with token-level value-guided decoding, while also reducing construction costs by over 30% compared to existing open-source PRMs. We also provide a thorough analysis and case study on its performance, transferability, and generalization capabilities. We provide our code on https://github.com/Lux0926/ASPRM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- DAPO : Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage-Based Policy OptimizationJiacai Liu, Chaojie Wang, Chris Yuhao Liu, Liang Zeng et al.NeurIPS 2025 · 9 citations
- Unlocking Token Rewards via Training-Free Reward AttributionWU Sitong, Haoru Tan, Bin Xia, Xichen Zhang et al.CVPR 2026
- TRAC: Teacher-Guided Token Reward with Adaptive Calibration for Robust Policy OptimizationSitong Wu, Haoru Tan, Xichen Zhang, Bin Xia et al.ACL 2026
- DesireKV: Decoupling Sensitivity and Importance for Reasoning-Aware KV Cache CompressionPengyu Cheng, Jiacheng Wang, Tianle Chen, Bei Liu et al.AAAI 2026
Builds on17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
Related papers
- Adversarial Training for Process Reward ModelsGurusha Juneja, Deepak Nathani, William WangICML 2026 · 2 citations
- From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time ScalingZhengyu Chen, Yudong Wang, Teng Xiao, Ruochen Zhou et al.AAAI 2026 · 2 citations
- Discriminative Policy Optimization for Token-Level Reward ModelsHongzhan Chen, Tao Yang, Shiping Gao, Ruijun Chen et al.ICML 2025
- Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical ReasoningJiuzhou Han, Wray L. Buntine, Ehsan ShareghiAAAI 2026 · 3 citations
- An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical ReasoningWei Sun, Qianlong Du, Fuwei Cui, Jiajun ZhangACL 2025 · 15 citations
