Offline Constrained Multi-Objective Reinforcement Learning via Pessimistic Dual Value Iteration
Runzhe Wu, Yufeng Zhang, Zhuoran Yang, Zhaoran Wang
摘要
In constrained multi-objective RL, the goal is to learn a policy that achieves the best performance specified by a multi-objective preference function under a constraint. We focus on the offline setting where the RL agent aims to learn the optimal policy from a given dataset. This scenario is common in real-world applications where interactions with the environment are expensive and the constraint violation is dangerous. For such a setting, we transform the original constrained problem into a primal-dual formulation, which is solved via dual gradient ascent. Moreover, we propose to combine such an approach with pessimism to overcome the uncertainty in offline data, which leads to our Pessimistic Dual Iteration (PEDI). We establish upper bounds on both the suboptimality and constraint violation for the policy learned by PEDI based on an arbitrary dataset, which proves that PEDI is provably sample efficient. We also specialize PEDI to the setting with linear function approximation. To the best of our knowledge, we propose the first provably efficient constrained multi-objective RL algorithm with offline data without any assumption on the coverage of the dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A Near-Optimal Primal-Dual Method for Off-Policy Learning in CMDPFan Chen, Junyu Zhang, Zaiwen WenNeurIPS 2022 · 被引用 15 次
- Bi-Level Offline Policy Optimization with Limited ExplorationWenzhuo ZhouNeurIPS 2023 · 被引用 6 次
- A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPsKihyuk Hong, Ambuj TewariICML 2024 · 被引用 5 次
- Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement LearningHonghao Wei, Xiyue Peng, Arnob Ghosh, Xin LiuNeurIPS 2024 · 被引用 4 次
- FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement LearningWoosung Kim, Jinho Lee, Jongmin Lee, Byung-Jun LeeNeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper14
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel 等NeurIPS 2020 · 被引用 406 次
相关 Paper
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 等ICLR 2022 · 被引用 84 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Scaling Pareto-Efficient Decision Making via Offline Multi-Objective RLBaiting Zhu, Meihua Dang, Aditya GroverICLR 2023 · 被引用 1 次
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 被引用 127 次
- POCE: Primal Policy Optimization with Conservative Estimation for Multi-constraint Offline Reinforcement LearningJiayi Guan, Li Shen, Ao Zhou, Lusong Li 等CVPR 2024
