Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning
Fan-Ming Luo, Tian Xu, Xingchen Cao, Yang Yu
摘要
Learning a precise dynamics model can be crucial for offline reinforcement learning, which, unfortunately, has been found to be quite challenging. Dynamics models that are learned by fitting historical transitions often struggle to generalize to unseen transitions. In this study, we identify a hidden but pivotal factor termed dynamics reward that remains consistent across transitions, offering a pathway to better generalization. Therefore, we propose the idea of reward-consistent dynamics models: any trajectory generated by the dynamics model should maximize the dynamics reward derived from the data. We implement this idea as the MOREC (Model-based Offline reinforcement learning with Reward Consistency) method, which can be seamlessly integrated into previous offline model-based reinforcement learning (MBRL) methods. MOREC learns a generalizable dynamics reward function from offline data, which is subsequently employed as a transition filter in any offline MBRL method: when generating transitions, the dynamics model generates a batch of transitions and selects the one with the highest dynamics reward value. On a synthetic task, we visualize that MOREC has a strong generalization ability and can surprisingly recover some distant unseen transitions. On 21 offline tasks in D4RL and NeoRL benchmarks, MOREC improves the previous state-of-the-art performance by a significant margin, i.e., 4.6% on D4RL tasks and 25.9% on NeoRL tasks. Notably, MOREC is the first method that can achieve above 95% online RL performance in 6 out of 12 D4RL tasks and 3 out of 9 NeoRL tasks. * CQL (Kumar et al., 2020) adds penalization to Q-values for the samples out of distribution; * TD3+BC (Fujimoto & Gu, 2021) incorporates a BC regularization term into the policy optimization objective; * EDAC (An et al., 2021) proposed to penalize based on the uncertainty degree of the Q-value. Model-based offline RL. * COMBO(Yu et al., 2021) which applies CQL in dyna-style enforces Q-values small on OOD samples; * RAMBO(Rigter et al., 2022) trains the dynamics model adversarially to minimize the value function without loss of accuracy on the transition prediction; * MOPO (Yu et al., 2020) learns a pessimistic value function from rewards penalized with the uncertainty of the dynamics model's prediction; * MOBILE (Sun et al., 2023) penalizes the rewards with uncertainty quantified by the inconsistency of Bellman estimations under an ensemble of learned dynamics models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory GenerationZhilong Zhang, Yihao Sun, Junyin Ye, Tian-Shuo Liu 等ICLR 2024 · 被引用 23 次
- KALM: Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model RolloutsJing-Cheng Pang, Si-Hang Yang, Kaiyuan Li, Jiaji Zhang 等NeurIPS 2024 · 被引用 12 次
- Grounded Answers for Multi-agent Decision-making Problem through Generative World ModelZeyang Liu, Xinrui Yang, Shiguang Sun, Long Qian 等NeurIPS 2024 · 被引用 10 次
- Offline Transition Modeling via Contrastive Energy LearningRuifeng Chen, Chengxing Jia, Zefang Huang, Tian-Shuo Liu 等ICML 2024 · 被引用 4 次
- Policy Learning from Tutorial Books via Understanding, Rehearsing and IntrospectingXiong-Hui Chen, Ziyan Wang, Yali Du, Shengyi Jiang 等NeurIPS 2024 · 被引用 4 次
它引用的顶会 Paper17
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 被引用 950 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
相关 Paper
- Model-Bellman Inconsistency for Model-based Offline Reinforcement LearningYihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin 等ICML 2023 · 被引用 61 次
- Model-based Offline Reinforcement Learning with Lower Expectile Q-LearningKwanyoung Park, Youngwoon LeeICLR 2025
- Optimistic Model Rollouts for Pessimistic Offline Policy OptimizationYuanzhao Zhai, Yiying Li, Zijian Gao, Xudong Gong 等AAAI 2024 · 被引用 4 次
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 被引用 18 次
- OCEAN-MBRL: Offline Conservative Exploration for Model-Based Offline Reinforcement LearningFan Wu, Rui Zhang, Qi Yi, Yunkai Gao 等AAAI 2024 · 被引用 4 次
