Trust Region Policy Optimization with Optimal Transport Discrepancies: Duality and Algorithm for Continuous Actions
Antonio Terpin, Nicolas Lanzetti, Batuhan Yardim, Florian Dörfler, Giorgia Ramponi
Abstract
Policy Optimization (PO) algorithms have been proven particularly suited to handle the high-dimensionality of real-world continuous control tasks. In this context, Trust Region Policy Optimization methods represent a popular approach to stabilize the policy updates. These usually rely on the Kullback-Leibler (KL) divergence to limit the change in the policy. The Wasserstein distance represents a natural alternative, in place of the KL divergence, to define trust regions or to regularize the objective function. However, state-of-the-art works either resort to its approximations or do not provide an algorithm for continuous state-action spaces, reducing the applicability of the method. In this paper, we explore optimal transport discrepancies (which include the Wasserstein distance) to define trust regions, and we propose a novel algorithm -Optimal Transport Trust Region Policy Optimization (OT-TRPO) -for continuous state-action spaces. We circumvent the infinite-dimensional optimization problem for PO by providing a one-dimensional dual reformulation for which strong duality holds. We then analytically derive the optimal policy update given the solution of the dual problem. This way, we bypass the computation of optimal transport costs and of optimal transport maps, which we implicitly characterize by solving the dual formulation. Finally, we provide an experimental evaluation of our approach across various control tasks. Our results show that optimal transport discrepancies can offer an advantage over state-of-theart approaches. * Equal contribution. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfe5f989-d2e3-49a1-bd5d-92d41754e6c6Cited by top-tier papers4
- Learning diffusion at lightspeedAntonio Terpin, Nicolas Lanzetti, Martín Gadea, Florian DörflerNeurIPS 2024 · 26 citations
- Pinet: Optimizing hard-constrained neural networks with orthogonal projection layersPanagiotis D. Grontas, Antonio Terpin, Efe C. Balta, Raffaello D'Andrea et al.ICLR 2026 · 22 citations
- BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement LearningYuan Li, Bo Wang, Yufei Gao, Yuqian Yao et al.ICML 2026 · 2 citations
- MRPO: Magnitude-Regularized Policy Optimization via L1 ConstraintsWei Han, Yuanxing Liu, Mingda Li, Ruiyu Xiao et al.ICML 2026
Builds on3
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Learning to Score Behaviors for Guided Policy OptimizationAldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Krzysztof Choromanski et al.ICML 2020 · 42 citations
- Efficient Wasserstein Natural Gradients for Reinforcement LearningTed Moskovitz, Michael Arbel, Ferenc Huszar, Arthur GrettonICLR 2021 · 23 citations
Related papers
- Differentiable Trust Region Layers for Deep Reinforcement LearningFabian Otto, Philipp Becker, Ngo Anh Vien, Hanna Carolin Maria Ziesche et al.ICLR 2021 · 23 citations
- Visual Transfer For Reinforcement Learning Via Wasserstein Domain ConfusionJosh Roy, George Dimitri KonidarisAAAI 2021 · 16 citations
- Policy Optimization for Continuous Reinforcement LearningHanyang Zhao, Wenpin Tang, David D. YaoNeurIPS 2023 · 47 citations
- Wasserstein Policy OptimizationDavid Pfau, Ian Davies, Diana L. Borsa, João Guilherme Madeira Araújo et al.ICML 2025
- Wasserstein Gradient Flows for Optimizing Gaussian Mixture PoliciesHanna Ziesche, Leonel RozoNeurIPS 2023 · 12 citations
