Wasserstein Policy Optimization
David Pfau, Ian Davies, Diana L. Borsa, João Guilherme Madeira Araújo, Brendan D. Tracey, Hado van Hasselt
摘要
We introduce Wasserstein Policy Optimization (WPO), an actor-critic algorithm for reinforcement learning in continuous action spaces. WPO can be derived as an approximation to Wasserstein gradient flow over the space of all policies projected into a finite-dimensional parameter space (e.g., the weights of a neural network), leading to a simple and completely general closed-form update. The resulting algorithm combines many properties of deterministic and classic policy gradient methods. Like deterministic policy gradients, it exploits knowledge of the gradient of the action-value function with respect to the action. Like classic policy gradients, it can be applied to stochastic policies with arbitrary distributions over actions -without using the reparameterization trick. We show results on the DeepMind Control Suite and a magnetic confinement fusion task which compare favorably with state-of-theart continuous control methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- SVGD as a kernelized Wasserstein gradient flow of the chi-squared divergenceSinho Chewi, Thibaut Le Gouic, Chen Lu, Tyler Maunu 等NeurIPS 2020 · 被引用 92 次
- Learning to Score Behaviors for Guided Policy OptimizationAldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Krzysztof Choromanski 等ICML 2020 · 被引用 42 次
- Wasserstein Unsupervised Reinforcement LearningShuncheng He, Yuhang Jiang, Hongchang Zhang, Jianzhun Shao 等AAAI 2022 · 被引用 30 次
- Wasserstein Quantum Monte Carlo: A Novel Approach for Solving the Quantum Many-Body Schrödinger EquationKirill Neklyudov, Jannes Nys, Luca A. Thiede, Juan Carrasquilla 等NeurIPS 2023 · 被引用 28 次
- Efficient Wasserstein Natural Gradients for Reinforcement LearningTed Moskovitz, Michael Arbel, Ferenc Huszar, Arthur GrettonICLR 2021 · 被引用 23 次
相关 Paper
- Mean Field Langevin Actor-Critic: Faster Convergence and Global Optimality beyond Lazy LearningKakei Yamamoto, Kazusato Oko, Zhuoran Yang, Taiji SuzukiICML 2024 · 被引用 2 次
- Flow Matching Policy GradientsDavid McAllister, Songwei Ge, Brent Yi, Chung Min Kim 等ICLR 2026 · 被引用 103 次
- Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions ControlAmarildo Likmeta, Matteo Sacco, Alberto Maria Metelli, Marcello RestelliAAAI 2023 · 被引用 7 次
- Promoting Stochasticity for Expressive Policies via a Simple and Efficient Regularization MethodQi Zhou, Yufei Kuang, Zherui Qiu, Houqiang Li 等NeurIPS 2020 · 被引用 9 次
- Parameter-Based Value FunctionsFrancesco Faccio, Louis Kirsch, Jürgen SchmidhuberICLR 2021 · 被引用 29 次
