Augmented Proximal Policy Optimization for Safe Reinforcement Learning
Juntao Dai, Jiaming Ji, Long Yang, Qian Zheng, Gang Pan
Abstract
Safe reinforcement learning considers practical scenarios that maximize the return while satisfying safety constraints. Current algorithms, which suffer from training oscillations or approximation errors, still struggle to update the policy efficiently with precise constraint satisfaction. In this article, we propose Augmented Proximal Policy Optimization (APPO), which augments the Lagrangian function of the primal constrained problem via attaching a quadratic deviation term. The constructed multiplier-penalty function dampens cost oscillation for stable convergence while being equivalent to the primal constrained problem to precisely control safety costs. APPO alternately updates the policy and the Lagrangian multiplier via solving the constructed augmented primal-dual problem, which can be easily implemented by any first-order optimizer. We apply our APPO methods in diverse safety-constrained tasks, setting a new state of the art compared with a comprehensive list of safe RL baselines. Extensive experiments verify the merits of our method in easy implementation, stable convergence, and precise cost control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext edd3241f-d73f-4de7-80c3-9c413dd2e97aCited by top-tier papers7
- SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained LearningBorong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei et al.NeurIPS 2025 · 84 citations
- SafeDreamer: Safe Reinforcement Learning with World ModelsWeidong Huang, Jiaming Ji, Chunhe Xia, Borong Zhang et al.ICLR 2024 · 46 citations
- VOCE: Variational Optimization with Conservative Estimation for Offline Safe Reinforcement LearningJiayi Guan, Guang Chen, Jiaming Ji, Long Yang et al.NeurIPS 2023 · 19 citations
- Safe Reinforcement Learning using Finite-Horizon Gradient-based EstimationJuntao Dai, Yaodong Yang, Qian Zheng, Gang PanICML 2024 · 3 citations
- e-COP : Episodic Constrained Optimization of PoliciesAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Sahil SinglaNeurIPS 2024 · 2 citations
Builds on4
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 238 citations
- IPO: Interior-Point Policy Optimization under ConstraintsYongshuai Liu, Jiaxin Ding, Xin LiuAAAI 2020 · 231 citations
Related papers
- CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement LearningAyoub Belouadah, Sylvain Kubler, YVES LE TRAONICML 2026
- Enhancing Efficiency of Safe Reinforcement Learning via Sample ManipulationShangding Gu, Laixi Shi, Yuhao Ding, Alois Knoll et al.NeurIPS 2024 · 14 citations
- Proactive Constrained Policy Optimization with Preemptive PenaltyNing Yang, Pengyu Wang, Guoqing Liu, Haifeng Zhang et al.AAAI 2026 · 1 citation
- Reward Penalties on Augmented States for Solving Richly Constrained RL EffectivelyHao Jiang, Tien Mai, Pradeep Varakantham, Huy HoangAAAI 2024 · 2 citations
- Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and ConstraintsYuhao Ding, Javad LavaeiAAAI 2023 · 32 citations
