Pareto Policy Pool for Model-based Offline Reinforcement Learning
Yijun Yang, Jing Jiang, Tianyi Zhou, Jie Ma, Yuhui Shi
摘要
Online reinforcement learning (RL) can suffer from poor exploration, sparse reward, insufficient data, and overhead caused by inefficient interactions between an immature policy and a complicated environment. Model-based offline RL instead trains an environment model using a dataset of pre-collected experiences so online RL methods can learn in an offline manner by solely interacting with the model. However, the uncertainty and accuracy of the environment model can drastically vary across different state-action pairs so the RL agent may achieve high model return but perform poorly in the true environment. Unlike previous works that need to carefully tune the trade-off between the model return and uncertainty in a single objective, we study a bi-objective formulation for model-based offline RL that aims at producing a pool of diverse policies on the Pareto front performing different levels of trade-offs, which provides the flexibility to select the best policy for each realistic environment from the pool. Our method, ''Pareto policy pool (P3)'', does not need to tune the trade-off weight but can produce policies allocated at different regions of the Pareto front. For this purpose, we develop an efficient algorithm that solves multiple bi-objective optimization problems with distinct constraints defined by reference vectors targeting diverse regions of the Pareto front. We theoretically prove that our algorithm can converge to the targeted regions. In order to obtain more Pareto optimal policies without linearly increasing the cost, we leverage the achieved policies as initialization to find more Pareto optimal policies in their neighborhoods. On the D4RL benchmark for offline RL, P3 substantially outperforms several recent baseline methods over multiple tasks, especially when the quality of pre-collected experiences is low.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper15
- RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2022 · 被引用 168 次
- Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-AvoidanceLisha Chen, Heshan Devaka Fernando, Yiming Ying, Tianyi ChenNeurIPS 2023 · 被引用 53 次
- Federated Multi-Objective LearningHaibo Yang, Zhuqing Liu, Jia Liu, Chaosheng Dong 等NeurIPS 2023 · 被引用 28 次
- One Risk to Rule Them All: A Risk-Sensitive Perspective on Model-Based Offline Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2023 · 被引用 26 次
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao 等CVPR 2024 · 被引用 23 次
相关 Paper
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Scaling Pareto-Efficient Decision Making via Offline Multi-Objective RLBaiting Zhu, Meihua Dang, Aditya GroverICLR 2023 · 被引用 1 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- OCEAN-MBRL: Offline Conservative Exploration for Model-Based Offline Reinforcement LearningFan Wu, Rui Zhang, Qi Yi, Yunkai Gao 等AAAI 2024 · 被引用 4 次
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement LearningChenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhi-Hong Deng 等ICLR 2022 · 被引用 173 次
