Sample Efficient Reinforcement Learning with REINFORCE
Junzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. Boyd
摘要
Policy gradient methods are among the most effective methods for large-scale reinforcement learning, and their empirical success has prompted several works that develop the foundation of their global convergence theory. However, prior works have either required exact gradients or state-action visitation measure based mini-batch stochastic gradients with a diverging batch size, which limit their applicability in practical scenarios. In this paper, we consider classical policy gradient methods that compute an approximate gradient with a single trajectory or a fixed size mini-batch of trajectories under soft-max parametrization and log-barrier regularization, along with the widely-used REINFORCE gradient estimation procedure. By controlling the number of "bad" episodes and resorting to the classical doubling trick, we establish an anytime sub-linear high probability regret bound as well as almost sure global convergence of the average regret with an asymptotically sub-linear rate. These provide the first set of global convergence and sample efficiency results for the well-known REINFORCE algorithm and contribute to a better understanding of its performance in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsZiniu Li, Tian Xu, Yushun Zhang, Zhihang Lin 等ICML 2024 · 被引用 165 次
- On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient MethodJunyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvári 等NeurIPS 2021 · 被引用 87 次
- GRACE: Loss-Resilient Real-Time Video through Neural CodecsYihua Cheng, Ziyi Zhang, Hanchen Li, Anton Arapin 等NSDI 2024 · 被引用 53 次
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM ReasoningKongcheng Zhang, Qi Yao, Shunyu Liu, Yingjie Wang 等NeurIPS 2025 · 被引用 45 次
- Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHFHan Shen, Zhuoran Yang, Tianyi ChenICML 2024 · 被引用 35 次
它引用的顶会 Paper13
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 被引用 304 次
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 被引用 270 次
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- Variational Policy Gradient Method for Reinforcement Learning with General UtilitiesJunyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvári 等NeurIPS 2020 · 被引用 170 次
相关 Paper
- Global Convergence of Policy Gradient in Average Reward MDPsNavdeep Kumar, Yashaswini Murthy, Itai Shufaro, Kfir Yehuda Levy 等ICLR 2025
- Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision ProcessesQinbo Bai, Washim Uddin Mondal, Vaneet AggarwalAAAI 2024 · 被引用 23 次
- Convergence and Optimality of Policy Gradient Methods in Weakly Smooth SettingsMatthew Shunshi Zhang, Murat A. Erdogdu, Animesh GargAAAI 2022 · 被引用 6 次
- REINFORCE Converges to Optimal Policies with Any Learning RateSamuel Robertson, Thang Chu, Bo Dai, Dale Schuurmans 等NeurIPS 2025 · 被引用 2 次
- Statistically Efficient Off-Policy Policy GradientsNathan Kallus, Masatoshi UeharaICML 2020 · 被引用 43 次
