Differentiable Trust Region Layers for Deep Reinforcement Learning
Fabian Otto, Philipp Becker, Ngo Anh Vien, Hanna Carolin Maria Ziesche, Gerhard Neumann
Abstract
Trust region methods are a popular tool in reinforcement learning as they yield robust policy updates in continuous and discrete action spaces. However, enforcing such trust regions in deep reinforcement learning is difficult. Hence, many approaches, such as Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), are based on approximations. Due to those approximations, they violate the constraints or fail to find the optimal solution within the trust region. Moreover, they are difficult to implement, often lack sufficient exploration, and have been shown to depend on seemingly unrelated implementation choices. In this work, we propose differentiable neural network layers to enforce trust regions for deep Gaussian policies via closed-form projections. Unlike existing methods, those layers formalize trust regions for each state individually and can complement existing reinforcement learning algorithms. We derive trust region projections based on the Kullback-Leibler divergence, the Wasserstein L2 distance, and the Frobenius norm for Gaussian distributions. We empirically demonstrate that those projection layers achieve similar or better results than existing methods while being almost agnostic to specific implementation choices. The code is available at https://git.io/Jthb0 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1a7690f-1c12-487d-a3d9-8c60c3bfed51Cited by top-tier papers17
- Trust Region Constrained Measure Transport in Path Space for Stochastic Optimal Control and InferenceDenis Blessing, Julius Berner, Lorenz Richter, Carles Domingo-Enrich et al.NeurIPS 2025 · 24 citations
- Acquiring Diverse Skills using Curriculum Reinforcement Learning with Mixture of ExpertsOnur Celik, Aleksandar Taranovic, Gerhard NeumannICML 2024 · 19 citations
- Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step ReturnsDong Tian, Onur Celik, Gerhard NeumannICLR 2026 · 18 citations
- Wasserstein Gradient Flows for Optimizing Gaussian Mixture PoliciesHanna Ziesche, Leonel RozoNeurIPS 2023 · 12 citations
- Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement LearningGe Li, Hongyi Zhou, Dominik Roth, Serge Thilges et al.ICLR 2024 · 11 citations
Builds on3
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous ControlH. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark et al.ICLR 2020 · 138 citations
- Learning to Score Behaviors for Guided Policy OptimizationAldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Krzysztof Choromanski et al.ICML 2020 · 42 citations
Related papers
- Trust Region Policy Optimization with Optimal Transport Discrepancies: Duality and Algorithm for Continuous ActionsAntonio Terpin, Nicolas Lanzetti, Batuhan Yardim, Florian Dörfler et al.NeurIPS 2022 · 14 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- Embedding Safety into RL: A New Take on Trust Region MethodsNikola Milosevic, Johannes Müller, Nico ScherfICML 2025
- Simple Policy OptimizationZhengpeng Xie, Qiang Zhang, Fan Yang, Marco Hutter et al.ICML 2025
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
