Conservative Bayesian Model-Based Value Expansion for Offline Policy Optimization
Jihwan Jeong, Xiaoyu Wang, Michael Gimelfarb, Hyunwoo Kim, Baher Abdulhai, Scott Sanner
Abstract
Offline reinforcement learning (RL) addresses the problem of learning a performant policy from a fixed batch of data collected by following some behavior policy. Model-based approaches are particularly appealing in the offline setting since they can extract more learning signals from the logged dataset by learning a model of the environment. However, the performance of existing model-based approaches falls short of model-free counterparts, due to the compounding of estimation errors in the learned model. Driven by this observation, we argue that it is critical for a model-based method to understand when to trust the model and when to rely on model-free estimates, and how to act conservatively w.r.t. both. To this end, we derive an elegant and simple methodology called conservative Bayesian model-based value expansion for offline policy optimization (CBOP), that trades off model-free and model-based estimates during the policy evaluation step according to their epistemic uncertainties, and facilitates conservatism by taking a lower bound on the Bayesian posterior value estimate. On the standard D4RL continuous control tasks, we find that our method significantly outperforms previous model-based approaches: e.g., MOPO by %, MOReL by % and COMBO by %. Further, CBOP achieves state-of-the-art performance on out of benchmark datasets while doing on par on the remaining datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e15c5cbc-2d96-412d-8cd6-3e52e68133c2Cited by top-tier papers12
- Model-Bellman Inconsistency for Model-based Offline Reinforcement LearningYihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin et al.ICML 2023 · 61 citations
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 19 citations
- Trust the Model Where It Trusts Itself - Model-Based Actor-Critic with Uncertainty-Aware Rollout AdaptionBernd Frauenknecht, Artur Eisele, Devdutt Subhasish, Friedrich Solowjow et al.ICML 2024 · 14 citations
- Scalable Offline Model-Based RL with Action ChunksKwanyoung Park, Seohong Park, Youngwoon Lee, Sergey LevineICLR 2026 · 12 citations
- Bayes Adaptive Monte Carlo Tree Search for Offline Model-based Reinforcement LearningJiayu Chen, Le Xu, Wen-Tse Chen, Jeff SchneiderICLR 2026 · 10 citations
Builds on15
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
Related papers
- OCEAN-MBRL: Offline Conservative Exploration for Model-Based Offline Reinforcement LearningFan Wu, Rui Zhang, Qi Yi, Yunkai Gao et al.AAAI 2024 · 4 citations
- Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian LensJihwan Jeong, Xiaoyu Wang, Jingmin Wang, Scott Sanner et al.ICML 2025
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
- VIPO: Value Function Inconsistency Penalized Offline Reinforcement LearningXuyang Chen, Keyu Yan, Guojian Wang, Lin ZhaoICML 2026 · 3 citations
- Model-based Offline Reinforcement Learning with Count-based ConservatismByeongchan Kim, Min-hwan OhICML 2023 · 19 citations
