VinePPO: Refining Credit Assignment in RL Training of LLMs
Amirhossein Kazemnejad, Milad Aghajohari, Eva Portelance, Alessandro Sordoni, Siva Reddy, Aaron C. Courville, Nicolas Le Roux
摘要
Large language models (LLMs) are increasingly applied to complex reasoning tasks that require executing several complex steps before receiving any reward. Properly assigning credit to these steps is essential for enhancing model performance. Proximal Policy Optimization (PPO), a common reinforcement learning (RL) algorithm used for LLM finetuning, employs value networks to tackle credit assignment. However, recent approaches achieve strong results without it, raising questions about the efficacy of value networks in practice. In this work, we systematically evaluate the efficacy of value networks and reveal their significant shortcomings in reasoning-heavy LLM tasks, showing that they often produce poor estimate of expected return and barely outperform a random baseline when comparing alternative steps. This motivates our key question: Can improved credit assignment enhance RL training for LLMs? To address this, we propose VinePPO, a straightforward approach that leverages the flexibility of language environments to compute unbiased Monte Carlo-based estimates. Our method consistently outperforms PPO and other baselines across MATH and GSM8K datasets in less wallclock time (up to 3.0x). Crucially, it achieves higher test accuracy for a given training accuracy, capturing more generalization signal per sample. These results emphasize the importance of accurate credit assignment in RL training of LLM. Code available at https://github.com/ McGill-NLP/VinePPO * Equal contribution † Equal advising
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Tree Search for LLM Agent Reinforcement LearningYuxiang Ji, Ziyu Ma, Yong Wang, Guanhua Chen 等ICLR 2026 · 被引用 71 次
- Prompt Curriculum Learning for Efficient LLM Post-TrainingZhaolin Gao, Joongwon Kim, Wen Sun, Thorsten Joachims 等ICLR 2026 · 被引用 44 次
- QuestA: Expanding Reasoning Capacity in LLMs via Question AugmentationJiazheng Li, Hongzhou Lin, Hong Lu, Kaiyue Wen 等ICLR 2026 · 被引用 42 次
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning ModelsRunze Liu, Jiakang Wang, Yuling Shi, Zhihui Xie 等ICLR 2026 · 被引用 13 次
- On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMsRosie Zhao, Anshul Shah, Xiaoyu Zhu, Xinke Deng 等ICML 2026 · 被引用 10 次
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
相关 Paper
- SSVPO: Effective Step-Level Credit Assignment for RL Training of Language ModelsYugu Li, Zehong Cao, Jianglin Qiao, Siyi HuICLR 2026
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning TasksTianyi Wang, Yixia Li, Long Li, Yibiao Chen 等ACL 2026 · 被引用 8 次
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang 等ICML 2026 · 被引用 22 次
- AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage MarginJian Xiong, Jingbo Zhou, Jingyong Ye, Qiang Huang 等ACL 2026 · 被引用 3 次
- SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and MoreMuye Huang, Lingling Zhang, Yifei Li, Yaqiang Wu 等CVPR 2026 · 被引用 7 次
