Predictable MDP Abstraction for Unsupervised Model-Based RL
Seohong Park, Sergey Levine
摘要
A key component of model-based reinforcement learning (RL) is a dynamics model that predicts the outcomes of actions. Errors in this predictive model can degrade the performance of modelbased controllers, and complex Markov decision processes (MDPs) can present exceptionally difficult prediction problems. To mitigate this issue, we propose predictable MDP abstraction (PMA): instead of training a predictive model on the original MDP, we train a model on a transformed MDP with a learned action space that only permits predictable, easy-to-model actions, while covering the original state-action space as much as possible. As a result, model learning becomes easier and more accurate, which allows robust, stable model-based planning or modelbased RL. This transformation is learned in an unsupervised manner, before any task is specified by the user. Downstream tasks can then be solved with model-based control in a zero-shot fashion, without additional environment interactions. We theoretically analyze PMA and empirically demonstrate that PMA leads to significant improvements over prior unsupervised modelbased RL approaches in a range of benchmark environments. Our code and videos are available at https://seohong.me/projects/pma/ Original MDP Predictable MDP
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement LearningRuijie Zheng, Xiyao Wang, Yanchao Sun, Shuang Ma 等NeurIPS 2023 · 被引用 89 次
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 被引用 83 次
- Can a MISL Fly? Analysis and Ingredients for Mutual Information Skill LearningChongyi Zheng, Jens Tuyls, Joanne Peng, Benjamin EysenbachICLR 2025
它引用的顶会 Paper25
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
相关 Paper
- Causal Dynamics Learning for Task-Independent State AbstractionZizhao Wang, Xuesu Xiao, Zifan Xu, Yuke Zhu 等ICML 2022 · 被引用 77 次
- Explainable Reinforcement Learning via Model TransformsMira Finkelstein, Nitsan Levy Schlot, Lucy Liu, Yoav Kolumbus 等NeurIPS 2022 · 被引用 18 次
- Model-based Reinforcement Learning for Parameterized Action SpacesRenhao Zhang, Haotian Fu, Yilin Miao, George KonidarisICML 2024 · 被引用 8 次
- Plan To Predict: Learning an Uncertainty-Foreseeing Model For Model-Based Reinforcement LearningZifan Wu, Chao Yu, Chen Chen, Jianye Hao 等NeurIPS 2022 · 被引用 28 次
- Planning with Abstract Learned Models While Learning Transferable SubtasksJohn Winder, Stephanie Milani, Matthew Landen, Erebus Oh 等AAAI 2020 · 被引用 10 次
