Trust the Model When It Is Confident: Masked Model-based Actor-Critic
Feiyang Pan, Jia He, Dandan Tu, Qing He
Abstract
It is a popular belief that model-based Reinforcement Learning (RL) is more sample efficient than model-free RL, but in practice, it is not always true due to overweighed model errors. In complex and noisy settings, model-based RL tends to have trouble using the model if it does not know when to trust the model. In this work, we find that better model usage can make a huge difference. We show theoretically that if the use of model-generated data is restricted to state-action pairs where the model error is small, the performance gap between model and real rollouts can be reduced. It motivates us to use model rollouts only when the model is confident about its predictions. We propose Masked Model-based Actor-Critic (M2AC), a novel policy optimization algorithm that maximizes a model-based lower-bound of the true value function. M2AC implements a masking mechanism based on the model's uncertainty to decide whether its prediction should be used or not. Consequently, the new algorithm tends to give robust policy improvements. Experiments on continuous control benchmarks demonstrate that M2AC has strong performance even when using long model rollouts in very noisy environments, and it significantly outperforms previous state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext accf7126-8a74-4458-b57d-594604a49e08Cited by top-tier papers20
- Continuous-time Model-based Reinforcement LearningÇagatay Yildiz, Markus Heinonen, Harri LähdesmäkiICML 2021 · 72 citations
- Revisiting Design Choices in Offline Model Based Reinforcement LearningCong Lu, Philip J. Ball, Jack Parker-Holder, Michael A. Osborne et al.ICLR 2022 · 65 citations
- MoCoDA: Model-based Counterfactual Data AugmentationSilviu Pitis, Elliot Creager, Ajay Mandlekar, Animesh GargNeurIPS 2022 · 60 citations
- Deconstructing the Inductive Biases of Hamiltonian Neural NetworksNate Gruver, Marc Anton Finzi, Samuel Don Stanton, Andrew Gordon WilsonICLR 2022 · 50 citations
- Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-CriticZhihai Wang, Jie Wang, Qi Zhou, Bin Li et al.AAAI 2022 · 38 citations
Builds on1
Related papers
- Trust the Model Where It Trusts Itself - Model-Based Actor-Critic with Uncertainty-Aware Rollout AdaptionBernd Frauenknecht, Artur Eisele, Devdutt Subhasish, Friedrich Solowjow et al.ICML 2024 · 14 citations
- Perceiving the Knowledge Boundary: Uncertainty-Guided Exploration and Imagination for World ModelsZhenxian Liu, Peixi Peng, Yangru Huang, Yonghong TianAAAI 2026
- On Rollouts in Model-Based Reinforcement LearningBernd Frauenknecht, Devdutt Subhasish, Friedrich Solowjow, Sebastian TrimpeICLR 2025 · 1 citation
- Model-Augmented Actor-Critic: Backpropagating through PathsIgnasi Clavera, Yao Fu, Pieter AbbeelICLR 2020 · 96 citations
- Deep Model-Based Reinforcement Learning via Estimated Uncertainty and Conservative Policy OptimizationQi Zhou, Houqiang Li, Jie WangAAAI 2020 · 17 citations
