POCE: Primal Policy Optimization with Conservative Estimation for Multi-constraint Offline Reinforcement Learning
Jiayi Guan, Li Shen, Ao Zhou, Lusong Li, Han Hu, Xiaodong He, Guang Chen, Changjun Jiang
摘要
Multi-constraint offline reinforcement learning (RL) promises to learn policies that satisfy both cumulative and state-wise costs from offline datasets. This arrangement provides an effective approach for the widespread application of RL in high-risk scenarios where both cumulative and state-wise costs need to be considered simultaneously. However, previously constrained offline RL algorithms are primarily designed to handle single-constraint problems related to cumulative cost, which faces challenges when addressing multi-constraint tasks that involve both cumulative and state-wise costs. In this work, we propose a novel Primal policy Optimization with Conservative Estimation algorithm (POCE) to address the problem of multi-constraint offline RL. Concretely, we reframe the objective of multi-constraint offline RL by introducing the concept of Maximum Markov Decision Processes (MMDP). Subsequently, we present a primal policy optimization algorithm to confront the multi-constraint problems, which improves the stability and convergence speed of model training. Furthermore, we propose a conditional Bellman operator to estimate cumulative and state-wise Qvalues, reducing the extrapolation error caused by out-ofdistribution (OOD) actions. Finally, extensive experiments demonstrate that the POCE algorithm achieves competitive performance across multiple experimental tasks, particularly outperforming baseline algorithms in terms of safety. Our code is available at github.POCE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement LearningYihang Yao, Zhepeng Cen, Wenhao Ding, Haohong Lin 等NeurIPS 2024 · 被引用 16 次
- An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement LearningQian Lin, Zongkai Liu, Danying Mo, Chao YuNeurIPS 2024 · 被引用 8 次
- CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement LearningAo Zhou, Jiayi Guan, Li Shen, Fan Lu 等NeurIPS 2025 · 被引用 1 次
- Direct Flow Q-LearningShicheng Cao, Jingrui Jia, Wenyu Li, Feng Duan 等ICML 2026
它引用的顶会 Paper30
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- Uncertainty-Based Offline Reinforcement Learning with Diversified Q-EnsembleGaon An, Seungyong Moon, Jang-Hyun Kim, Hyun Oh SongNeurIPS 2021 · 被引用 430 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
相关 Paper
- VOCE: Variational Optimization with Conservative Estimation for Offline Safe Reinforcement LearningJiayi Guan, Guang Chen, Jiaming Ji, Long Yang 等NeurIPS 2023 · 被引用 19 次
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 等ICLR 2022 · 被引用 84 次
- Off-Policy Primal-Dual Safe Reinforcement LearningZifan Wu, Bo Tang, Qian Lin, Chao Yu 等ICLR 2024 · 被引用 10 次
- C2IQL: Constraint-Conditioned Implicit Q-learning for Safe Offline Reinforcement LearningZifan Liu, Xinran Li, Jun ZhangICML 2025
- Online Optimization for Offline Safe Reinforcement LearningYassine Chemingui, Aryan Deshwal, Alan Fern, Thanh Nguyen-Tang 等NeurIPS 2025 · 被引用 3 次
