Efficient Wasserstein Natural Gradients for Reinforcement Learning
Ted Moskovitz, Michael Arbel, Ferenc Huszar, Arthur Gretton
Abstract
A novel optimization approach is proposed for application to policy gradient methods and evolution strategies for reinforcement learning (RL). The procedure uses a computationally efficient Wasserstein natural gradient (WNG) descent that takes advantage of the geometry induced by a Wasserstein penalty to speed optimization. This method follows the recent theme in RL of including a divergence penalty in the objective to establish a trust region. Experiments on challenging tasks demonstrate improvements in both computational cost and performance over advanced baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 770552ba-e4ca-45cd-9815-044801c859ccCited by top-tier papers16
- Confronting Reward Model Overoptimization with Constrained RLHFTed Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm et al.ICLR 2024 · 89 citations
- Minimum Description Length ControlTed Moskovitz, Ta-Chu Kao, Maneesh Sahani, Matt M. BotvinickICLR 2023 · 76 citations
- Tactical Optimism and Pessimism for Deep Reinforcement LearningTed Moskovitz, Jack Parker-Holder, Aldo Pacchiano, Michael Arbel et al.NeurIPS 2021 · 75 citations
- MICo: Improved representations via sampling-based state similarity for Markov decision processesPablo Samuel Castro, Tyler Kastner, Prakash Panangaden, Mark RowlandNeurIPS 2021 · 66 citations
- Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement LearningRishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, Marc G. BellemareICLR 2021 · 27 citations
Builds on2
Related papers
- Trust Region Policy Optimization with Optimal Transport Discrepancies: Duality and Algorithm for Continuous ActionsAntonio Terpin, Nicolas Lanzetti, Batuhan Yardim, Florian Dörfler et al.NeurIPS 2022 · 14 citations
- Mirror and Preconditioned Gradient Descent in Wasserstein SpaceClément Bonet, Théo Uscidda, Adam David, Pierre-Cyril Aubin-Frankowski et al.NeurIPS 2024 · 19 citations
- Wasserstein Policy OptimizationDavid Pfau, Ian Davies, Diana L. Borsa, João Guilherme Madeira Araújo et al.ICML 2025
- Semantic-aware Wasserstein Policy Regularization for Large Language Model AlignmentByeonghu Na, Hyungho Na, Yeongmin Kim, Suhyeon Jo et al.ICLR 2026 · 2 citations
- Tractable structured natural-gradient descent using local parameterizationsWu Lin, Frank Nielsen, Mohammad Emtiyaz Khan, Mark SchmidtICML 2021 · 36 citations
