Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions
Roland Stolz, Michael Eichelbeck, Matthias Althoff
Abstract
In reinforcement learning (RL), it is often advantageous to consider additional constraints on the action space to ensure safety or action relevance. Existing work on such action-constrained RL faces challenges regarding effective policy updates, computational efficiency, and predictable runtime. Recent work proposes to use truncated normal distributions for stochastic policy gradient methods. However, the computation of key characteristics, such as the entropy, log-probability, and their gradients, becomes intractable under complex constraints. Hence, prior work approximates these using the non-truncated distributions, which severely degrades performance. We argue that accurate estimation of these characteristics is crucial in the action-constrained RL setting, and propose efficient numerical approximations for them. We also provide an efficient sampling strategy for truncated policy distributions and validate our approach on three benchmark environments, which demonstrate significant performance improvements when using accurate estimations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e70eb29-33e0-4c0d-b025-0e08980cdb36Builds on5
- Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action MaskingRoland Stolz, Hanna Krasowski, Jakob Thumm, Michael Eichelbeck et al.NeurIPS 2024 · 30 citations
- FlowPG: Action-constrained Policy Gradient with Normalizing FlowsJanaka Chathuranga Brahmanage, Jiajing Ling, Akshat KumarNeurIPS 2023 · 16 citations
- Solving Online Threat Screening Games using Constrained Action Space Reinforcement LearningSanket Shah, Arunesh Sinha, Pradeep Varakantham, Andrew Perrault et al.AAAI 2020 · 14 citations
- Leveraging Constraint Violation Signals for Action Constrained Reinforcement LearningJanaka Chathuranga Brahmanage, Jiajing Ling, Akshat KumarAAAI 2025 · 2 citations
- Efficient Action-Constrained Reinforcement Learning via Acceptance-Rejection Method and Augmented MDPsWei Hung, Shao-Hua Sun, Ping-Chun HsiehICLR 2025
Related papers
- Tilted Quantile Gradient Updates for Quantile-Constrained Reinforcement LearningChenglin Li, Guangchun Ruan, Hua GengAAAI 2025
- Safe Reinforcement Learning using Finite-Horizon Gradient-based EstimationJuntao Dai, Yaodong Yang, Qian Zheng, Gang PanICML 2024 · 3 citations
- Learning with Safety Constraints: Sample Complexity of Reinforcement Learning for Constrained MDPsAria HasanzadeZonuzy, Archana Bura, Dileep M. Kalathil, Srinivas ShakkottaiAAAI 2021 · 46 citations
- Truncated Gaussian Policy for Debiased Continuous ControlGanghun Lee, Minji Kim, Minsu Lee, Byoung-Tak ZhangAAAI 2025 · 1 citation
- Trust Region-Based Safe Distributional Reinforcement Learning for Multiple ConstraintsDohyeong Kim, Kyungjae Lee, Songhwai OhNeurIPS 2023 · 26 citations
