Offline Model-based Adaptable Policy Learning
Xiong-Hui Chen, Yang Yu, Qingyang Li, Fan-Ming Luo, Zhiwei (Tony) Qin, Wenjie Shang, Jieping Ye
Abstract
In reinforcement learning, a promising direction to avoid online trial-and-error costs is learning from an offline dataset. Current offline reinforcement learning methods commonly learn in the policy space constrained to in-support regions by the offline dataset, in order to ensure the robustness of the outcome policies. Such constraints, however, also limit the potential of the outcome policies. In this paper, to release the potential of offline policy learning, we investigate the decision-making problems in out-of-support regions directly and propose offline Model-based Adaptable Policy LEarning (MAPLE). By this approach, instead of learning in in-support regions, we learn an adaptable policy that can adapt its behavior in out-of-support regions when deployed. We give a practical implementation of MAPLE via meta-learning techniques and ensemble model learning techniques. We conduct experiments on MuJoCo locomotion tasks with offline datasets. The results show that the proposed method can make robust decisions in out-of-support regions and achieve better performance than SOTA algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 369e6c21-7a28-4dbf-8c44-9668d183399fCited by top-tier papers17
- Model-Bellman Inconsistency for Model-based Offline Reinforcement LearningYihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin et al.ICML 2023 · 61 citations
- State Regularized Policy Optimization on Data with Dynamics ShiftZhenghai Xue, Qingpeng Cai, Shuchang Liu, Dong Zheng et al.NeurIPS 2023 · 30 citations
- Adversarial Counterfactual Environment Model LearningXiong-Hui Chen, Yang Yu, Zhengmao Zhu, Zhihua Yu et al.NeurIPS 2023 · 20 citations
- Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement LearningFan-Ming Luo, Tian Xu, Xingchen Cao, Yang YuICLR 2024 · 16 citations
- Conservative State Value Estimation for Offline Reinforcement LearningLiting Chen, Jie Yan, Zhengdao Shao, Lu Wang et al.NeurIPS 2023 · 15 citations
Builds on6
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
Related papers
- Supported Trust Region Optimization for Offline Reinforcement LearningYixiu Mao, Hongchang Zhang, Chen Chen, Yi Xu et al.ICML 2023 · 24 citations
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang et al.ICML 2022 · 78 citations
- Model-Based Offline Meta-Reinforcement Learning with RegularizationSen Lin, Jialin Wan, Tengyu Xu, Yingbin Liang et al.ICLR 2022 · 20 citations
- Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive LearningHaoqi Yuan, Zongqing LuICML 2022 · 53 citations
- Meta-Reinforcement Learning Based on Self-Supervised Task Representation LearningMingyang Wang, Zhenshan Bing, Xiangtong Yao, Shuai Wang et al.AAAI 2023 · 22 citations
