Mismatched No More: Joint Model-Policy Optimization for Model-Based RL
Benjamin Eysenbach, Alexander Khazatsky, Sergey Levine, Ruslan Salakhutdinov
Abstract
Many model-based reinforcement learning (RL) methods follow a similar template: fit a model to previously observed data, and then use data from that model for RL or planning. However, models that achieve better training performance (e.g., lower MSE) are not necessarily better for control: an RL agent may seek out the small fraction of states where an accurate model makes mistakes, or it might act in ways that do not expose the errors of an inaccurate model. As noted in prior work, there is an objective mismatch: models are useful if they yield good policies, but they are trained to maximize their accuracy, rather than the performance of the policies that result from them. In this work, we propose a single objective for jointly training the model and the policy, such that updates to either component increase a lower bound on expected return. To the best of our knowledge, this is the first lower bound for model-based RL that holds globally and can be efficiently estimated in continuous settings; it is the only lower bound that mends the objective mismatch problem. A version of this bound becomes tight under certain assumptions. Optimizing this bound resembles a GAN: a classifier distinguishes between real and fake transitions, the model is updated to produce transitions that look realistic, and the policy is updated to avoid states where the model predictions are unrealistic. Numerical simulations demonstrate that optimizing this bound yields reward maximizing policies and yields dynamics that (perhaps surprisingly) can aid in exploration. We also show that a deep RL algorithm loosely based on our lower bound can achieve performance competitive with prior model-based methods, and better performance on certain hard exploration tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eecf15a7-2363-4ba2-916d-bae59565f441Cited by top-tier papers25
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- Generative Learning for Financial Time Series with Irregular and Scale-Invariant PatternsHongbin Huang, Minghua Chen, Xiao QiaoICLR 2024 · 51 citations
- Preference-grounded Token-level Guidance for Language Model Fine-tuningShentao Yang, Shujian Zhang, Congying Xia, Yihao Feng et al.NeurIPS 2023 · 39 citations
- Beyond OOD State Actions: Supported Cross-Domain Offline Reinforcement LearningJinxin Liu, Ziqi Zhang, Zhenyu Wei, Zifeng Zhuang et al.AAAI 2024 · 30 citations
- Maximize to Explore: One Objective Function Fusing Estimation, Planning, and ExplorationZhihan Liu, Miao Lu, Wei Xiong, Han Zhong et al.NeurIPS 2023 · 30 citations
Builds on11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
- A Game Theoretic Framework for Model Based Reinforcement LearningAravind Rajeswaran, Igor Mordatch, Vikash KumarICML 2020 · 137 citations
Related papers
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
- Model-based Policy Optimization with Unsupervised Model AdaptationJian Shen, Han Zhao, Weinan Zhang, Yong YuNeurIPS 2020 · 33 citations
- Sample-Efficient Iterative Lower Bound Optimization of Deep Reactive Policies for Planning in Continuous MDPsSiow Meng Low, Akshat Kumar, Scott SannerAAAI 2022 · 3 citations
- The Virtues of Laziness in Model-based RL: A Unified Objective and AlgorithmsAnirudh Vemula, Yuda Song, Aarti Singh, Drew Bagnell et al.ICML 2023 · 15 citations
- Model-Based Offline Reinforcement Learning with Local MisspecificationKefan Dong, Yannis Flet-Berliac, Allen Nie, Emma BrunskillAAAI 2023 · 6 citations
