Constraints Penalized Q-learning for Safe Offline Reinforcement Learning
Haoran Xu, Xianyuan Zhan, Xiangyu Zhu
Abstract
We study the problem of safe offline reinforcement learning (RL), the goal is to learn a policy that maximizes long-term reward while satisfying safety constraints given only offline data, without further interaction with the environment. This problem is more appealing for real world RL applications, in which data collection is costly or dangerous. Enforcing constraint satisfaction is non-trivial, especially in offline settings, as there is a potential large discrepancy between the policy distribution and the data distribution, causing errors in estimating the value of safety constraints. We show that naïve approaches that combine techniques from safe RL and offline RL can only learn sub-optimal solutions. We thus develop a simple yet effective algorithm, Constraints Penalized Q-Learning (CPQ), to solve the problem. Our method admits the use of data generated by mixed behavior policies. We present a theoretical analysis and demonstrate empirically that our approach can learn robustly across a variety of benchmark control tasks, outperforming several baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aad430fa-3b43-48cd-b4f6-6a966ff879dcCited by top-tier papers54
- DeepThermal: Combustion Optimization for Thermal Power Generating Units Using Offline Reinforcement LearningXianyuan Zhan, Haoran Xu, Yue Zhang, Xiangyu Zhu et al.AAAI 2022 · 96 citations
- A Policy-Guided Imitation Approach for Offline Reinforcement LearningHaoran Xu, Li Jiang, Jianxiong Li, Xianyuan ZhanNeurIPS 2022 · 86 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
- Constrained Decision Transformer for Offline Safe Reinforcement LearningZuxin Liu, Zijian Guo, Yihang Yao, Zhepeng Cen et al.ICML 2023 · 82 citations
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li et al.NeurIPS 2022 · 81 citations
Builds on7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- DeepThermal: Combustion Optimization for Thermal Power Generating Units Using Offline Reinforcement LearningXianyuan Zhan, Haoran Xu, Yue Zhang, Xiangyu Zhu et al.AAAI 2022 · 96 citations
Related papers
- Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement LearningPrajwal Koirala, Zhanhong Jiang, Soumik Sarkar, Cody H. FlemingICLR 2025
- C2IQL: Constraint-Conditioned Implicit Q-learning for Safe Offline Reinforcement LearningZifan Liu, Xinran Li, Jun ZhangICML 2025
- Online Optimization for Offline Safe Reinforcement LearningYassine Chemingui, Aryan Deshwal, Alan Fern, Thanh Nguyen-Tang et al.NeurIPS 2025 · 3 citations
- Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RLJunyu Guo, Zhi Zheng, Donghao Ying, Ming Jin et al.NeurIPS 2025 · 2 citations
- Offline Safe Reinforcement Learning Using Trajectory ClassificationZe Gong, Akshat Kumar, Pradeep VarakanthamAAAI 2025 · 6 citations
