Post-Training with Policy Gradients: Optimality and the Base Model Barrier
Alireza Mousavi-Hosseini, Murat Erdogdu
摘要
We study post-training linear autoregressive models with outcome and process rewards. Given a context , the model must predict the response , a sequence of length that satisfies a standard margin assumption extended to sequences. We prove that on test samples where the base model achieves a non-trivial likelihood , a variant of policy gradient (PG) can achieve likelihood with an essentially minimax optimal number of reward queries . However, a barrier arises for going beyond the support of the base model. We prove that the overall expected error after post-training with outcome rewards is governed by a property of the base model we call the Likelihood Quantile (LQ), and that variants of PG, while minimax optimal, may require a number of reward queries exponential in to go beyond this support, regardless of the pre-training algorithm. To overcome this barrier, we study post-training with a process reward model, and demonstrate how PG variants in this setting avoid the curse of dimensionality in via dependence on a token-level LQ. Along the way, we prove that under the margin condition, SGD with adaptive learning rate (LR) achieves a near optimal test error for statistical learning, and PG with adaptive LR achieves a near optimal number of mistakes for online learning while being computationally efficient whenever possible, both of which may be of independent interest.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient MethodsYanli Liu, Kaiqing Zhang, Tamer Basar, Wotao YinNeurIPS 2020 · 被引用 128 次
- On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient MethodJunyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvári 等NeurIPS 2021 · 被引用 87 次
- Reasoning with Sampling: Your Base Model is Smarter Than You ThinkAayush Karan, Yilun DuICLR 2026 · 被引用 87 次
- The Coverage Principle: How Pre-Training Enables Post-TrainingFan Chen, Audrey Huang, Noah Golowich, Sadhika Malladi 等ICLR 2026 · 被引用 28 次
相关 Paper
- Robust Policy Gradient against Strong Data CorruptionXuezhou Zhang, Yiding Chen, Xiaojin Zhu, Wen SunICML 2021 · 被引用 43 次
- MLE-Guided Parameter Search for Task Loss Minimization in Neural Sequence ModelingSean Welleck, Kyunghyun ChoAAAI 2021 · 被引用 8 次
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
- Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage ProbabilityWhiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul SungNeurIPS 2022 · 被引用 12 次
- Ordering-based Conditions for Global Convergence of Policy Gradient MethodsJincheng Mei, Bo Dai, Alekh Agarwal, Mohammad Ghavamzadeh 等NeurIPS 2023 · 被引用 4 次
