Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning
Fan-Ming Luo, Tian Xu, Xingchen Cao, Yang Yu
Abstract
Learning a precise dynamics model can be crucial for offline reinforcement learning, which, unfortunately, has been found to be quite challenging. Dynamics models that are learned by fitting historical transitions often struggle to generalize to unseen transitions. In this study, we identify a hidden but pivotal factor termed dynamics reward that remains consistent across transitions, offering a pathway to better generalization. Therefore, we propose the idea of reward-consistent dynamics models: any trajectory generated by the dynamics model should maximize the dynamics reward derived from the data. We implement this idea as the MOREC (Model-based Offline reinforcement learning with Reward Consistency) method, which can be seamlessly integrated into previous offline model-based reinforcement learning (MBRL) methods. MOREC learns a generalizable dynamics reward function from offline data, which is subsequently employed as a transition filter in any offline MBRL method: when generating transitions, the dynamics model generates a batch of transitions and selects the one with the highest dynamics reward value. On a synthetic task, we visualize that MOREC has a strong generalization ability and can surprisingly recover some distant unseen transitions. On 21 offline tasks in D4RL and NeoRL benchmarks, MOREC improves the previous state-of-the-art performance by a significant margin, i.e., 4.6% on D4RL tasks and 25.9% on NeoRL tasks. Notably, MOREC is the first method that can achieve above 95% online RL performance in 6 out of 12 D4RL tasks and 3 out of 9 NeoRL tasks. * CQL (Kumar et al., 2020) adds penalization to Q-values for the samples out of distribution; * TD3+BC (Fujimoto & Gu, 2021) incorporates a BC regularization term into the policy optimization objective; * EDAC (An et al., 2021) proposed to penalize based on the uncertainty degree of the Q-value. Model-based offline RL. * COMBO(Yu et al., 2021) which applies CQL in dyna-style enforces Q-values small on OOD samples; * RAMBO(Rigter et al., 2022) trains the dynamics model adversarially to minimize the value function without loss of accuracy on the transition prediction; * MOPO (Yu et al., 2020) learns a pessimistic value function from rewards penalized with the uncertainty of the dynamics model's prediction; * MOBILE (Sun et al., 2023) penalizes the rewards with uncertainty quantified by the inconsistency of Bellman estimations under an ensemble of learned dynamics models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Flow to Better: Offline Preference-based Reinforcement Learning via Preferred Trajectory GenerationZhilong Zhang, Yihao Sun, Junyin Ye, Tian-Shuo Liu et al.ICLR 2024 · 23 citations
- KALM: Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model RolloutsJing-Cheng Pang, Si-Hang Yang, Kaiyuan Li, Jiaji Zhang et al.NeurIPS 2024 · 12 citations
- Grounded Answers for Multi-agent Decision-making Problem through Generative World ModelZeyang Liu, Xinrui Yang, Shiguang Sun, Long Qian et al.NeurIPS 2024 · 10 citations
- Offline Transition Modeling via Contrastive Energy LearningRuifeng Chen, Chengxing Jia, Zefang Huang, Tian-Shuo Liu et al.ICML 2024 · 4 citations
- Policy Learning from Tutorial Books via Understanding, Rehearsing and IntrospectingXiong-Hui Chen, Ziyan Wang, Yali Du, Shengyi Jiang et al.NeurIPS 2024 · 4 citations
Builds on17
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
Related papers
- Model-Bellman Inconsistency for Model-based Offline Reinforcement LearningYihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin et al.ICML 2023 · 61 citations
- Model-based Offline Reinforcement Learning with Lower Expectile Q-LearningKwanyoung Park, Youngwoon LeeICLR 2025
- Optimistic Model Rollouts for Pessimistic Offline Policy OptimizationYuanzhao Zhai, Yiying Li, Zijian Gao, Xudong Gong et al.AAAI 2024 · 4 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
- OCEAN-MBRL: Offline Conservative Exploration for Model-Based Offline Reinforcement LearningFan Wu, Rui Zhang, Qi Yi, Yunkai Gao et al.AAAI 2024 · 4 citations
