Mismatched No More: Joint Model-Policy Optimization for Model-Based RL
Benjamin Eysenbach, Alexander Khazatsky, Sergey Levine, Ruslan Salakhutdinov
摘要
Many model-based reinforcement learning (RL) methods follow a similar template: fit a model to previously observed data, and then use data from that model for RL or planning. However, models that achieve better training performance (e.g., lower MSE) are not necessarily better for control: an RL agent may seek out the small fraction of states where an accurate model makes mistakes, or it might act in ways that do not expose the errors of an inaccurate model. As noted in prior work, there is an objective mismatch: models are useful if they yield good policies, but they are trained to maximize their accuracy, rather than the performance of the policies that result from them. In this work, we propose a single objective for jointly training the model and the policy, such that updates to either component increase a lower bound on expected return. To the best of our knowledge, this is the first lower bound for model-based RL that holds globally and can be efficiently estimated in continuous settings; it is the only lower bound that mends the objective mismatch problem. A version of this bound becomes tight under certain assumptions. Optimizing this bound resembles a GAN: a classifier distinguishes between real and fake transitions, the model is updated to produce transitions that look realistic, and the policy is updated to avoid states where the model predictions are unrealistic. Numerical simulations demonstrate that optimizing this bound yields reward maximizing policies and yields dynamics that (perhaps surprisingly) can aid in exploration. We also show that a deep RL algorithm loosely based on our lower bound can achieve performance competitive with prior model-based methods, and better performance on certain hard exploration tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 被引用 1,115 次
- Generative Learning for Financial Time Series with Irregular and Scale-Invariant PatternsHongbin Huang, Minghua Chen, Xiao QiaoICLR 2024 · 被引用 51 次
- Preference-grounded Token-level Guidance for Language Model Fine-tuningShentao Yang, Shujian Zhang, Congying Xia, Yihao Feng 等NeurIPS 2023 · 被引用 39 次
- Beyond OOD State Actions: Supported Cross-Domain Offline Reinforcement LearningJinxin Liu, Ziqi Zhang, Zhenyu Wei, Zifeng Zhuang 等AAAI 2024 · 被引用 30 次
- Maximize to Explore: One Objective Function Fusing Estimation, Planning, and ExplorationZhihan Liu, Miao Lu, Wei Xiong, Han Zhong 等NeurIPS 2023 · 被引用 30 次
它引用的顶会 Paper11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran 等NeurIPS 2021 · 被引用 549 次
- A Game Theoretic Framework for Model Based Reinforcement LearningAravind Rajeswaran, Igor Mordatch, Vikash KumarICML 2020 · 被引用 137 次
相关 Paper
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 被引用 18 次
- Model-based Policy Optimization with Unsupervised Model AdaptationJian Shen, Han Zhao, Weinan Zhang, Yong YuNeurIPS 2020 · 被引用 33 次
- Sample-Efficient Iterative Lower Bound Optimization of Deep Reactive Policies for Planning in Continuous MDPsSiow Meng Low, Akshat Kumar, Scott SannerAAAI 2022 · 被引用 3 次
- The Virtues of Laziness in Model-based RL: A Unified Objective and AlgorithmsAnirudh Vemula, Yuda Song, Aarti Singh, Drew Bagnell 等ICML 2023 · 被引用 15 次
- Model-Based Offline Reinforcement Learning with Local MisspecificationKefan Dong, Yannis Flet-Berliac, Allen Nie, Emma BrunskillAAAI 2023 · 被引用 6 次
