Relative Entropy Pathwise Policy Optimization
Claas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman, Pieter Abbeel, Radu Grosu, Eric Eaton, Amir-massoud Farahmand, Igor Gilitschenski
摘要
Score-function based methods for policy learning, such as REINFORCE and PPO, have delivered strong results in game-playing and robotics, yet their high variance often undermines training stability. Improving a policy through state-action value functions, for example by differentiating Q with regard to the policy, alleviates the variance issues. However, this requires an accurate action-conditioned value function, which is notoriously hard to learn without relying on replay buffers for reusing past off-policy data. We present Relative Entropy Pathwise Policy Optimization, an algorithm that trains Q-value models purely from on-policy trajectories, unlocking the use of Q function derivatives to compute policy updates in the context of on-policy learning. We show how to combine stochastic policies for exploration with constrained updates for stable training, and evaluate important architectural components that stabilize value function learning. This results in an efficient on-policy algorithm that combines the stability of Q-based policy gradients with the simplicity and minimal memory footprint of standard on-policy learning. Compared to state-of-the-art on two standard GPU-parallelized benchmarks, REPPO provides strong empirical performance at superior sample efficiency, wall-clock time, memory footprint, and hyperparameter robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper48
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm 等ICLR 2021 · 被引用 399 次
相关 Paper
- Model-Augmented Actor-Critic: Backpropagating through PathsIgnasi Clavera, Yao Fu, Pieter AbbeelICLR 2020 · 被引用 96 次
- Generalized Proximal Policy Optimization with Sample ReuseJames Queeney, Yannis Paschalidis, Christos G. CassandrasNeurIPS 2021 · 被引用 80 次
- PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement LearningShunpeng Yang, Ben Liu, Hua ChenICLR 2026 · 被引用 6 次
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous ControlH. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark 等ICLR 2020 · 被引用 138 次
- Ratio-Variance Regularized Policy OptimizationYu Luo, Shuo Han, Yihan Hu, Lei Lv 等ICML 2026
