Constrained Policy Optimization with Explicit Behavior Density For Offline Reinforcement Learning
Jing Zhang, Chi Zhang, Wenjia Wang, Bingyi Jing
Abstract
Due to the inability to interact with the environment, offline reinforcement learning (RL) methods face the challenge of estimating the Out-of-Distribution (OOD) points. Existing methods for addressing this issue either control policy to exclude the OOD action or make the function pessimistic. However, these methods can be overly conservative or fail to identify OOD areas accurately. To overcome this problem, we propose a Constrained Policy optimization with Explicit Behavior density (CPED) method that utilizes a flow-GAN model to explicitly estimate the density of behavior policy. By estimating the explicit density, CPED can accurately identify the safe region and enable optimization within the region, resulting in less conservative learning policies. We further provide theoretical results for both the flow-GAN estimator and performance guarantee for CPED by showing that CPED can find the optimal -function value. Empirically, CPED outperforms existing alternatives on various standard offline reinforcement learning tasks, yielding higher expected returns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- GenPO: Generative Diffusion Models Meet On-Policy Reinforcement LearningShutong Ding, Ke Hu, Shan Zhong, Haoyang Luo et al.NeurIPS 2025 · 22 citations
- Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency modelJing Zhang, Linjiajie Fang, Kexin Shi, Wenjia Wang et al.NeurIPS 2024 · 14 citations
- Iteratively Refined Behavior Regularization for Offline Reinforcement LearningYi Ma, Jianye Hao, Xiaohan Hu, Yan Zheng et al.NeurIPS 2024 · 11 citations
- ReFORM: Reflected Flows for On-support Offline RL via Noise ManipulationSongyuan Zhang, Oswin So, H. M. Sabbir Ahmad, Eric Yang Yu et al.ICLR 2026 · 5 citations
- Adaptive Neighborhood-Constrained Q Learning for Offline Reinforcement LearningYixiu Mao, Yun Qu, Qi (Cheems) Wang, Xiangyang JiNeurIPS 2025 · 3 citations
Builds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
Related papers
- Flow Actor-Critic for Offline Reinforcement LearningJongseong Chae, Jongeui Park, Yongjae Shin, Gyeongmin Kim et al.ICLR 2026 · 7 citations
- Supported Policy Optimization for Offline Reinforcement LearningJialong Wu, Haixu Wu, Zihan Qiu, Jianmin Wang et al.NeurIPS 2022 · 113 citations
- Learning from Sparse Offline Datasets via Conservative Density EstimationZhepeng Cen, Zuxin Liu, Zitong Wang, Yihang Yao et al.ICLR 2024 · 12 citations
- Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization StrategiesRunze Yan, Xun Shen, Akifumi Wachi, Sebastien Gros et al.NeurIPS 2025 · 7 citations
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 127 citations
