Post-Training with Policy Gradients: Optimality and the Base Model Barrier
Alireza Mousavi-Hosseini, Murat Erdogdu
Abstract
We study post-training linear autoregressive models with outcome and process rewards. Given a context , the model must predict the response , a sequence of length that satisfies a standard margin assumption extended to sequences. We prove that on test samples where the base model achieves a non-trivial likelihood , a variant of policy gradient (PG) can achieve likelihood with an essentially minimax optimal number of reward queries . However, a barrier arises for going beyond the support of the base model. We prove that the overall expected error after post-training with outcome rewards is governed by a property of the base model we call the Likelihood Quantile (LQ), and that variants of PG, while minimax optimal, may require a number of reward queries exponential in to go beyond this support, regardless of the pre-training algorithm. To overcome this barrier, we study post-training with a process reward model, and demonstrate how PG variants in this setting avoid the curse of dimensionality in via dependence on a token-level LQ. Along the way, we prove that under the margin condition, SGD with adaptive learning rate (LR) achieves a near optimal test error for statistical learning, and PG with adaptive LR achieves a near optimal number of mistakes for online learning while being computationally efficient whenever possible, both of which may be of independent interest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ce2b3bc-ab25-4d26-ae7a-92c044f3936fBuilds on10
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient MethodsYanli Liu, Kaiqing Zhang, Tamer Basar, Wotao YinNeurIPS 2020 · 128 citations
- On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient MethodJunyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvári et al.NeurIPS 2021 · 87 citations
- Reasoning with Sampling: Your Base Model is Smarter Than You ThinkAayush Karan, Yilun DuICLR 2026 · 87 citations
- The Coverage Principle: How Pre-Training Enables Post-TrainingFan Chen, Audrey Huang, Noah Golowich, Sadhika Malladi et al.ICLR 2026 · 28 citations
Related papers
- Robust Policy Gradient against Strong Data CorruptionXuezhou Zhang, Yiding Chen, Xiaojin Zhu, Wen SunICML 2021 · 43 citations
- MLE-Guided Parameter Search for Task Loss Minimization in Neural Sequence ModelingSean Welleck, Kyunghyun ChoAAAI 2021 · 8 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
- Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage ProbabilityWhiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul SungNeurIPS 2022 · 12 citations
- Ordering-based Conditions for Global Convergence of Policy Gradient MethodsJincheng Mei, Bo Dai, Alekh Agarwal, Mohammad Ghavamzadeh et al.NeurIPS 2023 · 4 citations
