Plan To Predict: Learning an Uncertainty-Foreseeing Model For Model-Based Reinforcement Learning
Zifan Wu, Chao Yu, Chen Chen, Jianye Hao, Hankz Hankui Zhuo
Abstract
In Model-based Reinforcement Learning (MBRL), model learning is critical since an inaccurate model can bias policy learning via generating misleading samples. However, learning an accurate model can be difficult since the policy is continually updated and the induced distribution over visited states used for model learning shifts accordingly. Prior methods alleviate this issue by quantifying the uncertainty of model-generated samples. However, these methods only quantify the uncertainty passively after the samples were generated, rather than foreseeing the uncertainty before model trajectories fall into those highly uncertain regions. The resulting low-quality samples can induce unstable learning targets and hinder the optimization of the policy. Moreover, while being learned to minimize one-step prediction errors, the model is generally used to predict for multiple steps, leading to a mismatch between the objectives of model learning and model usage. To this end, we propose Plan To Predict (P2P), an MBRL framework that treats the model rollout process as a sequential decision making problem by reversely considering the model as a decision maker and the current policy as the dynamics. In this way, the model can quickly adapt to the current policy and foresee the multi-step future uncertainty when generating trajectories. Theoretically, we show that the performance of P2P can be guaranteed by approximately optimizing a lower bound of the true environment return. Empirical results demonstrate that P2P achieves state-of-the-art performance on several challenging benchmark tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- How to Fine-tune the Model: Unified Model Shift and Model Bias Policy OptimizationHai Zhang, Hang Yu, Junqiao Zhao, Di Zhang et al.NeurIPS 2023 · 16 citations
- Trust the Model Where It Trusts Itself - Model-Based Actor-Critic with Uncertainty-Aware Rollout AdaptionBernd Frauenknecht, Artur Eisele, Devdutt Subhasish, Friedrich Solowjow et al.ICML 2024 · 14 citations
- Prioritized Model Experience ReplayMuxi Tao, jiangtao wen, Yuxing HanICML 2026
- Robust Reinforcement Learning in Finance: Modeling Market Impact with Elliptic Uncertainty SetsShaocong Ma, Heng HuangNeurIPS 2025
- Relative Policy-Transition Optimization for Fast Policy TransferJiawei Xu, Cheng Zhou, Yizheng Zhang, Baoxiang Wang et al.AAAI 2024
Builds on8
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Error Bounds of Imitating Policies and EnvironmentsTian Xu, Ziniu Li, Yang YuNeurIPS 2020 · 141 citations
- Goal-Aware Prediction: Learning to Model What MattersSuraj Nair, Silvio Savarese, Chelsea FinnICML 2020 · 71 citations
- Trust the Model When It Is Confident: Masked Model-based Actor-CriticFeiyang Pan, Jia He, Dandan Tu, Qing HeNeurIPS 2020 · 65 citations
Related papers
- On Rollouts in Model-Based Reinforcement LearningBernd Frauenknecht, Devdutt Subhasish, Friedrich Solowjow, Sebastian TrimpeICLR 2025 · 1 citation
- Bidirectional Model-based Policy OptimizationHang Lai, Jian Shen, Weinan Zhang, Yong YuICML 2020 · 66 citations
- COPlanner: Plan to Roll Out Conservatively but to Explore Optimistically for Model-Based RLXiyao Wang, Ruijie Zheng, Yanchao Sun, Ruonan Jia et al.ICLR 2024 · 19 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
- Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian LensJihwan Jeong, Xiaoyu Wang, Jingmin Wang, Scott Sanner et al.ICML 2025
