V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu, Dhruva Tirumala, Nicolas Heess, Dan Belov
Abstract
Some of the most successful applications of deep reinforcement learning to challenging domains in discrete and continuous control have used policy gradient methods in the on-policy setting. However, policy gradients can suffer from large variance that may limit performance, and in practice require carefully tuned entropy regularization to prevent policy collapse. As an alternative to policy gradient algorithms, we introduce V-MPO, an on-policy adaptation of Maximum a Posteriori Policy Optimization (MPO) that performs policy iteration based on a learned state-value function. We show that V-MPO surpasses previously reported scores for both the Atari-57 and DMLab-30 benchmark suites in the multi-task setting, and does so reliably without importance weighting, entropy regularization, or population-based tuning of hyperparameters. On individual DMLab and Atari levels, the proposed algorithm can achieve scores that are substantially higher than has previously been reported. V-MPO is also applicable to problems with high-dimensional, continuous action spaces, which we demonstrate in the context of learning to control simulated humanoids with 22 degrees of freedom from full state observations and 56 degrees of freedom from pixel observations, as well as example OpenAI Gym tasks where V-MPO achieves substantially higher asymptotic scores than previously reported.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd45bad1-9bb2-4641-9706-67da555ab9e7Cited by top-tier papers19
- Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement LearningYiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng et al.NeurIPS 2021 · 133 citations
- My Body is a Cage: the Role of Morphology in Graph-Based Incompatible ControlVitaly Kurin, Maximilian Igl, Tim Rocktäschel, Wendelin Boehmer et al.ICLR 2021 · 105 citations
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert et al.ICML 2020 · 84 citations
- DrM: Mastering Visual Reinforcement Learning through Dormant Ratio MinimizationGuowei Xu, Ruijie Zheng, Yongyuan Liang, Xiyao Wang et al.ICLR 2024 · 53 citations
- Efficient Transformers in Reinforcement Learning using Actor-Learner DistillationEmilio Parisotto, Ruslan SalakhutdinovICLR 2021 · 51 citations
Related papers
- Robust Reinforcement Learning for Continuous Control with Model MisspecificationDaniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki et al.ICLR 2020 · 138 citations
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman et al.ICLR 2026 · 6 citations
- A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor RepresentationScott Fujimoto, David Meger, Doina PrecupICML 2021 · 17 citations
- Delay-Adapted Policy Optimization and Improved Regret for Adversarial MDP with Delayed Bandit FeedbackTal Lancewicki, Aviv Rosenberg, Dmitry SotnikovICML 2023 · 6 citations
- Mean-Variance Policy Iteration for Risk-Averse Reinforcement LearningShangtong Zhang, Bo Liu, Shimon WhitesonAAAI 2021 · 44 citations
