Wasserstein Policy Optimization
David Pfau, Ian Davies, Diana L. Borsa, João Guilherme Madeira Araújo, Brendan D. Tracey, Hado van Hasselt
Abstract
We introduce Wasserstein Policy Optimization (WPO), an actor-critic algorithm for reinforcement learning in continuous action spaces. WPO can be derived as an approximation to Wasserstein gradient flow over the space of all policies projected into a finite-dimensional parameter space (e.g., the weights of a neural network), leading to a simple and completely general closed-form update. The resulting algorithm combines many properties of deterministic and classic policy gradient methods. Like deterministic policy gradients, it exploits knowledge of the gradient of the action-value function with respect to the action. Like classic policy gradients, it can be applied to stochastic policies with arbitrary distributions over actions -without using the reparameterization trick. We show results on the DeepMind Control Suite and a magnetic confinement fusion task which compare favorably with state-of-theart continuous control methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b18c7ae-fff9-4295-9384-4ef7a7f5c989Cited by top-tier papers1
Ask how each one uses itBuilds on6
- SVGD as a kernelized Wasserstein gradient flow of the chi-squared divergenceSinho Chewi, Thibaut Le Gouic, Chen Lu, Tyler Maunu et al.NeurIPS 2020 · 92 citations
- Learning to Score Behaviors for Guided Policy OptimizationAldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Krzysztof Choromanski et al.ICML 2020 · 42 citations
- Wasserstein Unsupervised Reinforcement LearningShuncheng He, Yuhang Jiang, Hongchang Zhang, Jianzhun Shao et al.AAAI 2022 · 30 citations
- Wasserstein Quantum Monte Carlo: A Novel Approach for Solving the Quantum Many-Body Schrödinger EquationKirill Neklyudov, Jannes Nys, Luca A. Thiede, Juan Carrasquilla et al.NeurIPS 2023 · 28 citations
- Efficient Wasserstein Natural Gradients for Reinforcement LearningTed Moskovitz, Michael Arbel, Ferenc Huszar, Arthur GrettonICLR 2021 · 23 citations
Related papers
- Mean Field Langevin Actor-Critic: Faster Convergence and Global Optimality beyond Lazy LearningKakei Yamamoto, Kazusato Oko, Zhuoran Yang, Taiji SuzukiICML 2024 · 2 citations
- Flow Matching Policy GradientsDavid McAllister, Songwei Ge, Brent Yi, Chung Min Kim et al.ICLR 2026 · 103 citations
- Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions ControlAmarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello RestelliAAAI 2023 · 7 citations
- Promoting Stochasticity for Expressive Policies via a Simple and Efficient Regularization MethodQi Zhou, Yufei Kuang, Zherui Qiu, Houqiang Li et al.NeurIPS 2020 · 9 citations
- Parameter-Based Value FunctionsFrancesco Faccio, Louis Kirsch, Jürgen SchmidhuberICLR 2021 · 29 citations
