AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
Yuliang Liu, Junjie Lu, Chaofeng Qu, Zhaoling Chen, Zefan Cai, Jason Klein Liu, Chonghan Liu, Yunhui Xia, Li Zhao, Jiang Bian, Chuheng Zhang, Wei Shen, Zhouhan Lin
摘要
Current approaches for training Process Reward Models (PRMs) often involve decomposing responses into multiple reasoning steps using rulebased techniques, such as using predefined placeholder tokens or setting the reasoning step's length to a fixed size. These approaches overlook the fact that certain words don't usually indicate true decision points. To address this, we propose AdaptiveStep, a method that divides reasoning steps based on the model's confidence in predicting the next word, offering more information on decision-making at each step, improving downstream tasks like reward model training. Moreover, our method requires no manual annotation. Experiments with AdaptiveStep-trained PRMs in mathematical reasoning and code generation show that the outcome PRM achieves state-of-the-art Best-of-N performance, surpassing greedy search strategy with token-level value-guided decoding, while also reducing construction costs by over 30% compared to existing open-source PRMs. We also provide a thorough analysis and case study on its performance, transferability, and generalization capabilities. We provide our code on https://github.com/Lux0926/ASPRM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DAPO : Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage-Based Policy OptimizationJiacai Liu, Chaojie Wang, Chris Yuhao Liu, Liang Zeng 等NeurIPS 2025 · 被引用 9 次
- Unlocking Token Rewards via Training-Free Reward AttributionWU Sitong, Haoru Tan, Bin Xia, Xichen Zhang 等CVPR 2026
- TRAC: Teacher-Guided Token Reward with Adaptive Calibration for Robust Policy OptimizationSitong Wu, Haoru Tan, Xichen Zhang, Bin Xia 等ACL 2026
- DesireKV: Decoupling Sensitivity and Importance for Reasoning-Aware KV Cache CompressionPengyu Cheng, Jiacheng Wang, Tianle Chen, Bei Liu 等AAAI 2026
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
相关 Paper
- Adversarial Training for Process Reward ModelsGurusha Juneja, Deepak Nathani, William WangICML 2026 · 被引用 2 次
- From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time ScalingZhengyu Chen, Yudong Wang, Teng Xiao, Ruochen Zhou 等AAAI 2026 · 被引用 2 次
- Discriminative Policy Optimization for Token-Level Reward ModelsHongzhan Chen, Tao Yang, Shiping Gao, Ruijun Chen 等ICML 2025
- Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical ReasoningJiuzhou Han, Wray L. Buntine, Ehsan ShareghiAAAI 2026 · 被引用 3 次
- An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical ReasoningWei Sun, Qianlong Du, Fuwei Cui, Jiajun ZhangACL 2025 · 被引用 15 次
