Predictable MDP Abstraction for Unsupervised Model-Based RL
Seohong Park, Sergey Levine
Abstract
A key component of model-based reinforcement learning (RL) is a dynamics model that predicts the outcomes of actions. Errors in this predictive model can degrade the performance of modelbased controllers, and complex Markov decision processes (MDPs) can present exceptionally difficult prediction problems. To mitigate this issue, we propose predictable MDP abstraction (PMA): instead of training a predictive model on the original MDP, we train a model on a transformed MDP with a learned action space that only permits predictable, easy-to-model actions, while covering the original state-action space as much as possible. As a result, model learning becomes easier and more accurate, which allows robust, stable model-based planning or modelbased RL. This transformation is learned in an unsupervised manner, before any task is specified by the user. Downstream tasks can then be solved with model-based control in a zero-shot fashion, without additional environment interactions. We theoretically analyze PMA and empirically demonstrate that PMA leads to significant improvements over prior unsupervised modelbased RL approaches in a range of benchmark environments. Our code and videos are available at https://seohong.me/projects/pma/ Original MDP Predictable MDP
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ccf9f2f-98e6-4c0c-87c3-d88aeba71bd9Cited by top-tier papers3
- TACO: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement LearningRuijie Zheng, Xiyao Wang, Yanchao Sun, Shuang Ma et al.NeurIPS 2023 · 89 citations
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 83 citations
- Can a MISL Fly? Analysis and Ingredients for Mutual Information Skill LearningChongyi Zheng, Jens Tuyls, Joanne Peng, Benjamin EysenbachICLR 2025
Builds on25
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
Related papers
- Causal Dynamics Learning for Task-Independent State AbstractionZizhao Wang, Xuesu Xiao, Zifan Xu, Yuke Zhu et al.ICML 2022 · 77 citations
- Explainable Reinforcement Learning via Model TransformsMira Finkelstein, Nitsan Levy Schlot, Lucy Liu, Yoav Kolumbus et al.NeurIPS 2022 · 18 citations
- Model-based Reinforcement Learning for Parameterized Action SpacesRenhao Zhang, Haotian Fu, Yilin Miao, George KonidarisICML 2024 · 8 citations
- Plan To Predict: Learning an Uncertainty-Foreseeing Model For Model-Based Reinforcement LearningZifan Wu, Chao Yu, Chen Chen, Jianye Hao et al.NeurIPS 2022 · 28 citations
- Planning with Abstract Learned Models While Learning Transferable SubtasksJohn Winder, Stephanie Milani, Matthew Landen, Erebus Oh et al.AAAI 2020 · 10 citations
