Deep Model-Based Reinforcement Learning via Estimated Uncertainty and Conservative Policy Optimization
Qi Zhou, Houqiang Li, Jie Wang
摘要
Model-based reinforcement learning algorithms tend to achieve higher sample efficiency than model-free methods. However, due to the inevitable errors of learned models, model-based methods struggle to achieve the same asymptotic performance as model-free methods. In this paper, We propose a Policy Optimization method with Model-Based Uncertainty (POMBU)—a novel model-based approach—that can effectively improve the asymptotic performance using the uncertainty in Q-values. We derive an upper bound of the uncertainty, based on which we can approximate the uncertainty accurately and efficiently for model-based methods. We further propose an uncertainty-aware policy optimization algorithm that optimizes the policy conservatively to encourage performance improvement with high probability. This can significantly alleviate the overfitting of policy to inaccurate models. Experiments show POMBU can outperform existing state-of-the-art policy optimization algorithms in terms of sample efficiency and asymptotic performance. Moreover, the experiments demonstrate the excellent robustness of POMBU compared to previous model-based approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-CriticZhihai Wang, Jie Wang, Qi Zhou, Bin Li 等AAAI 2022 · 被引用 38 次
- Learning to Stop Cut Generation for Efficient Mixed-Integer Linear ProgrammingHaotian Ling, Zhihai Wang, Jie WangAAAI 2024 · 被引用 14 次
- Universal Value-Function UncertaintiesMoritz Akiya Zanger, Max Weltevrede, Yaniv Oren, Pascal R. van der Vaart 等ICLR 2026 · 被引用 1 次
- Epistemic Monte Carlo Tree SearchYaniv Oren, Viliam Vadocz, Matthijs T. J. Spaan, Wendelin BoehmerICLR 2025
- IL-SOAR : Imitation Learning with Soft Optimistic Actor cRiticStefano Viel, Luca Viano, Volkan CevherICML 2025
相关 Paper
- Model-Augmented Actor-Critic: Backpropagating through PathsIgnasi Clavera, Yao Fu, Pieter AbbeelICLR 2020 · 被引用 96 次
- Bidirectional Model-based Policy OptimizationHang Lai, Jian Shen, Weinan Zhang, Yong YuICML 2020 · 被引用 66 次
- Making Better Decision by Directly Planning in Continuous ControlJinhua Zhu, Yue Wang, Lijun Wu, Tao Qin 等ICLR 2023
- Trust the Model When It Is Confident: Masked Model-based Actor-CriticFeiyang Pan, Jia He, Dandan Tu, Qing HeNeurIPS 2020 · 被引用 65 次
- Constrained Policy Optimization via Bayesian World ModelsYarden As, Ilnura Usmanova, Sebastian Curi, Andreas KrauseICLR 2022 · 被引用 73 次
