Reduced Policy Optimization for Continuous Control with Hard Constraints
Shutong Ding, Jingya Wang, Yali Du, Ye Shi
摘要
Recent advances in constrained reinforcement learning (RL) have endowed reinforcement learning with certain safety guarantees. However, deploying existing constrained RL algorithms in continuous control tasks with general hard constraints remains challenging, particularly in those situations with non-convex hard constraints. Inspired by the generalized reduced gradient (GRG) algorithm, a classical constrained optimization technique, we propose a reduced policy optimization (RPO) algorithm that combines RL with GRG to address general hard constraints. RPO partitions actions into basic actions and nonbasic actions following the GRG method and outputs the basic actions via a policy network. Subsequently, RPO calculates the nonbasic actions by solving equations based on equality constraints using the obtained basic actions. The policy network is then updated by implicitly differentiating nonbasic actions with respect to basic actions. Additionally, we introduce an action projection procedure based on the reduced gradient and apply a modified Lagrangian relaxation technique to ensure inequality constraints are satisfied. To the best of our knowledge, RPO is the first attempt that introduces GRG to RL as a way of efficiently handling both equality and inequality hard constraints. It is worth noting that there is currently a lack of RL environments with complex hard constraints, which motivates us to develop three new benchmarks: two robotics manipulation tasks and a smart grid operation control task. With these benchmarks, RPO achieves better performance than previous constrained RL algorithms in terms of both cumulative reward and constraint violation. We believe RPO, along with the new benchmarks, will open up new opportunities for applying RL to real-world problems with complex constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Differentiation Through Black-Box Quadratic Programming SolversConnor W. Magoon, Fengyu Yang, Noam Aigerman, Shahar Z. KovalskyNeurIPS 2025 · 被引用 14 次
- Autoregressive Policy Optimization for Constrained Allocation TasksDavid Winkel, Niklas Strauß, Maximilian Bernhard, Zongyue Li 等NeurIPS 2024 · 被引用 2 次
- Diffusion-based learning framework for Constrained Nonconvex Optimization with Weighted Bootstrapped RefinementShutong Ding, Yimiao Zhou, Ke Hu, Xi Yao 等ICML 2026 · 被引用 1 次
- Gauge Flow Matching: Efficient Constrained Generative Modeling over General Convex Set and BeyondXinpeng Li, Enming Liang, Minghua ChenICLR 2026
- Safety-Polarized and Prioritized Reinforcement LearningKe Fan, Jinpeng Zhang, Xuefeng Zhang, Yunze Wu 等ICML 2025
它引用的顶会 Paper10
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 被引用 403 次
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 被引用 238 次
- IPO: Interior-Point Policy Optimization under ConstraintsYongshuai Liu, Jiaxin Ding, Xin LiuAAAI 2020 · 被引用 231 次
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 被引用 190 次
相关 Paper
- Constrained Reinforcement Learning Under Model MismatchZhongchang Sun, Sihong He, Fei Miao, Shaofeng ZouICML 2024 · 被引用 12 次
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 被引用 171 次
- Gradient-Adaptive Pareto Optimization for Constrained Reinforcement LearningZixian Zhou, Mengda Huang, Feiyang Pan, Jia He 等AAAI 2023 · 被引用 11 次
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 被引用 58 次
- Embedding Safety into RL: A New Take on Trust Region MethodsNikola Milosevic, Johannes Müller, Nico ScherfICML 2025
