VinePPO: Refining Credit Assignment in RL Training of LLMs
Amirhossein Kazemnejad, Milad Aghajohari, Eva Portelance, Alessandro Sordoni, Siva Reddy, Aaron C. Courville, Nicolas Le Roux
Abstract
Large language models (LLMs) are increasingly applied to complex reasoning tasks that require executing several complex steps before receiving any reward. Properly assigning credit to these steps is essential for enhancing model performance. Proximal Policy Optimization (PPO), a common reinforcement learning (RL) algorithm used for LLM finetuning, employs value networks to tackle credit assignment. However, recent approaches achieve strong results without it, raising questions about the efficacy of value networks in practice. In this work, we systematically evaluate the efficacy of value networks and reveal their significant shortcomings in reasoning-heavy LLM tasks, showing that they often produce poor estimate of expected return and barely outperform a random baseline when comparing alternative steps. This motivates our key question: Can improved credit assignment enhance RL training for LLMs? To address this, we propose VinePPO, a straightforward approach that leverages the flexibility of language environments to compute unbiased Monte Carlo-based estimates. Our method consistently outperforms PPO and other baselines across MATH and GSM8K datasets in less wallclock time (up to 3.0x). Crucially, it achieves higher test accuracy for a given training accuracy, capturing more generalization signal per sample. These results emphasize the importance of accurate credit assignment in RL training of LLM. Code available at https://github.com/ McGill-NLP/VinePPO * Equal contribution † Equal advising
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 865e6e30-ccf1-488b-a720-ee0ac6a45f0dCited by top-tier papers19
- Tree Search for LLM Agent Reinforcement LearningYuxiang Ji, Ziyu Ma, Yong Wang, Guanhua Chen et al.ICLR 2026 · 71 citations
- Prompt Curriculum Learning for Efficient LLM Post-TrainingZhaolin Gao, Joongwon Kim, Wen Sun, Thorsten Joachims et al.ICLR 2026 · 44 citations
- QuestA: Expanding Reasoning Capacity in LLMs via Question AugmentationJiazheng Li, Hongzhou Lin, Hong Lu, Kaiyue Wen et al.ICLR 2026 · 42 citations
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning ModelsRunze Liu, Jiakang Wang, Yuling Shi, Zhihui Xie et al.ICLR 2026 · 13 citations
- On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMsRosie Zhao, Anshul Shah, Xiaoyu Zhu, Xinke Deng et al.ICML 2026 · 10 citations
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
Related papers
- SSVPO: Effective Step-Level Credit Assignment for RL Training of Language ModelsYugu Li, Zehong Cao, Jianglin Qiao, Siyi HuICLR 2026
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning TasksTianyi Wang, Yixia Li, Long Li, Yibiao Chen et al.ACL 2026 · 8 citations
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang et al.ICML 2026 · 22 citations
- AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage MarginJian Xiong, Jingbo Zhou, Jingyong Ye, Qiang Huang et al.ACL 2026 · 3 citations
- SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and MoreMuye Huang, Lingling Zhang, Yifei Li, Yaqiang Wu et al.CVPR 2026 · 7 citations
