Accelerating RL for LLM Reasoning with Optimal Advantage Regression
Kianté Brantley, Mingyu Chen, Zhaolin Gao, Jason D. Lee, Wen Sun, Wenhao Zhan, Xuezhou Zhang
摘要
Reinforcement learning (RL) has emerged as a powerful tool for fine-tuning large language models (LLMs) to improve complex reasoning abilities. However, stateof-the-art policy optimization methods often suffer from high computational overhead and memory consumption, primarily due to the need for multiple generations per prompt and the reliance on critic networks or advantage estimates of the current policy. In this paper, we propose A ⋆ -PO, a novel two-stage policy optimization framework that directly approximates the optimal advantage function and enables efficient training of LLMs for reasoning tasks. In the first stage, we leverage offline sampling from a reference policy to estimate the optimal value function V ⋆ , eliminating the need for costly online value estimation. In the second stage, we perform on-policy updates using a simple least-squares regression loss with only a single generation per prompt. Theoretically, we establish performance guarantees and prove that the KL-regularized RL objective can be optimized without requiring complex exploration strategies. Empirically, A ⋆ -PO achieves competitive performance across a wide range of mathematical reasoning benchmarks, while reducing training time by up to 2× and peak memory usage by over 30% compared to PPO, GRPO, and REBEL [Gao et al., 2024a]. Implementation of A ⋆ -PO can be found at https://github.com/ZhaolinGao/A-PO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Prompt Curriculum Learning for Efficient LLM Post-TrainingZhaolin Gao, Joongwon Kim, Wen Sun, Thorsten Joachims 等ICLR 2026 · 被引用 44 次
- Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-TrainingBrian R. Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain 等NeurIPS 2025 · 被引用 34 次
- Single-stream Policy OptimizationZhongwen Xu, Zihan DingICLR 2026 · 被引用 29 次
- Stable and Efficient Single-Rollout RL for Multimodal ReasoningRui Liu, Dian Yu, Lei Ke, Haolin Liu 等CVPR 2026 · 被引用 13 次
- Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its FriendsChaorui Yao, Yanxi Chen, Yuchang Sun, Yushuo Chen 等ICLR 2026 · 被引用 13 次
它引用的顶会 Paper18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
相关 Paper
- AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage MarginJian Xiong, Jingbo Zhou, Jingyong Ye, Qiang Huang 等ACL 2026 · 被引用 3 次
- DAPO : Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage-Based Policy OptimizationJiacai Liu, Chaojie Wang, Chris Yuhao Liu, Liang Zeng 等NeurIPS 2025 · 被引用 9 次
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng 等ICLR 2026 · 被引用 4 次
- GPG: A Simple and Strong Reinforcement Learning Baseline for Model ReasoningXiangxiang Chu, Hailang Huang, Xiao Zhang, Fei Wei 等ICLR 2026 · 被引用 168 次
- GPO: Learning from Critical Steps to Improve LLM ReasoningJiahao Yu, Zelei Cheng, Xian Wu, Xinyu XingNeurIPS 2025 · 被引用 10 次
