Efficient Action-Constrained Reinforcement Learning via Acceptance-Rejection Method and Augmented MDPs
Wei Hung, Shao-Hua Sun, Ping-Chun Hsieh
Abstract
Action-constrained reinforcement learning (ACRL) is a generic framework for learning control policies with zero action constraint violation, which is required by various safety-critical and resource-constrained applications. The existing ACRL methods can typically achieve favorable constraint satisfaction but at the cost of either a high computational burden incurred by the quadratic programs (QP) or increased architectural complexity due to the use of sophisticated generative models. In this paper, we propose a generic and computationally efficient framework that can adapt a standard unconstrained RL method to ACRL through two modifications: (i) To enforce the action constraints, we leverage the classic acceptancerejection method, where we treat the unconstrained policy as the proposal distribution and derive a modified policy with feasible actions. (ii) To improve the acceptance rate of the proposal distribution, we construct an augmented two-objective Markov decision process (MDP), which includes additional self-loop state transitions and a penalty signal for the rejected actions. This augmented MDP incentivizes the learned policy to stay close to the feasible action sets. Through extensive experiments in both robot control and resource allocation domains, we demonstrate that the proposed framework enjoys faster training progress, better constraint satisfaction, and a lower action inference time simultaneously than the state-of-the-art ACRL methods. We have made the source code publicly available * to encourage further research in this direction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- SAD-Flower: Flow Matching for Safe, Admissible, and Dynamically Consistent PlanningTzu-Yuan Huang, Armin Lederer, Dai-Jie Wu, Xiaobing Dai et al.ICML 2026 · 3 citations
- Cross-Domain Policy Optimization via Bellman Consistency and Hybrid CriticsMing-Hong Chen, Kuan-Chen Pan, You-De Huang, Xi Liu et al.ICLR 2026 · 1 citation
- Improving Stochastic Action-Constrained Reinforcement Learning via Truncated DistributionsRoland Stolz, Michael Eichelbeck, Matthias AlthoffAAAI 2026 · 1 citation
- Action-Constrained Imitation LearningChia-Han Yeh, Tse-Sheng Nan, Risto Vuorio, Wei Hung et al.ICML 2025
Builds on6
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 238 citations
- Safe Reinforcement Learning by Imagining the Near FutureGarrett Thomas, Yuping Luo, Tengyu MaNeurIPS 2021 · 118 citations
- Accelerating Quadratic Optimization with Reinforcement LearningJeffrey Ichnowski, Paras Jain, Bartolomeo Stellato, Goran Banjac et al.NeurIPS 2021 · 62 citations
- Towards Safe Reinforcement Learning with a Safety Editor PolicyHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2022 · 50 citations
- FlowPG: Action-constrained Policy Gradient with Normalizing FlowsJanaka Chathuranga Brahmanage, Jiajing Ling, Akshat KumarNeurIPS 2023 · 16 citations
Related papers
- Leveraging Constraint Violation Signals for Action Constrained Reinforcement LearningJanaka Chathuranga Brahmanage, Jiajing Ling, Akshat KumarAAAI 2025 · 2 citations
- Robust Inverse Constrained Reinforcement Learning under Model MisspecificationSheng Xu, Guiliang LiuICML 2024 · 7 citations
- Conservative and Adaptive Penalty for Model-Based Safe Reinforcement LearningYecheng Jason Ma, Andrew Shen, Osbert Bastani, Dinesh JayaramanAAAI 2022 · 32 citations
- CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement LearningAo Zhou, Jiayi Guan, Li Shen, Fan Lu et al.NeurIPS 2025 · 1 citation
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 58 citations
