GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
Alexander Havrilla, Sharath Chandra Raparthy, Christoforos Nalmpantis, Jane Dwivedi-Yu, Maksym Zhuravinskyi, Eric Hambro, Roberta Raileanu
摘要
State-of-the-art language models can exhibit impressive reasoning refinement capabilities on math, science or coding tasks. However, recent work demonstrates that even the best models struggle to identify when and where to refine without access to external feedback. Outcome-based Reward Models (ORMs), trained to predict correctness of the final answer indicating when to refine, offer one convenient solution. However, when used to indicate where to refine, we find that ORMs tend to be overly-pessimistic when used to assess intermediate reasoning steps, resulting in excessive refinement of valid solutions. Process Based Reward Models (PRMs), trained to predict correctness of intermediate steps indicating where to refine, have been used to improve LLM reasoning ability via rejection sampling or reinforcement learning (RL) fine-tuning. But they are expensive to train, requiring extensive human annotations. In this paper, we propose Stepwise ORMs (SORMs) which are trained, only on synthetic data, to approximate the expected future reward of the optimal policy or V ⋆ . More specifically, SORMs are trained to predict the correctness of the final answer when sampling the current policy many times (rather than only once as in the case of ORMs). Our experiments show that SORMs can more accurately detect incorrect reasoning steps compared to ORMs, thus improving downstream accuracy when doing refinements. We then train global refinement models, which take only the question and a draft solution as input and predict a corrected solution, and local refinement models which also take as input a critique indicating the location of the first reasoning error. We generate training data for both models synthetically by reusing data used to train the SORM. We find combining global and local refinements, using the ORM as a reranker, significantly outperforms either one individually, as well as a best of three sample baseline. With this strategy we can improve the accuracy of a LLaMA-2 13B model (already fine-tuned with RL) on GSM8K from 53% to 65% when greedily sampled.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper44
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Recursive Introspection: Teaching Language Model Agents How to Self-ImproveYuxiao Qu, Tianjun Zhang, Naman Garg, Aviral KumarNeurIPS 2024 · 被引用 218 次
- Darwin Gödel Machine: Open-Ended Evolution of Self-Improving AgentsJenny Zhang, Shengran Hu, Cong Lu, Robert Tjarko Lange 等ICLR 2026 · 被引用 101 次
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement LearningQianyue Hao, Sibo Li, Jian Yuan, Yong LiICLR 2026 · 被引用 20 次
- Adaptive Batch-Wise Sample Scheduling for Direct Preference OptimizationZixuan Huang, Yikun Ban, Lean Fu, Xiaojie Li 等NeurIPS 2025 · 被引用 14 次
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
相关 Paper
- Discriminative Policy Optimization for Token-Level Reward ModelsHongzhan Chen, Tao Yang, Shiping Gao, Ruijun Chen 等ICML 2025
- MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement LearningZhiheng Xi, Yuhui Wang, Yiwen Ding, Guanyu Li 等AAAI 2026
- Generative Verifiers: Reward Modeling as Next-Token PredictionLunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi 等ICLR 2025
- Rewarding Progress: Scaling Automated Process Verifiers for LLM ReasoningAmrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng 等ICLR 2025
- Linking Process to Outcome: Conditional Reward Modeling for LLM ReasoningZheng Zhang, Ziwei Shan, Kaitao Song, Yexin Li 等ICLR 2026 · 被引用 16 次
