Conservative Bayesian Model-Based Value Expansion for Offline Policy Optimization
Jihwan Jeong, Xiaoyu Wang, Michael Gimelfarb, Hyunwoo Kim, Baher Abdulhai, Scott Sanner
摘要
Offline reinforcement learning (RL) addresses the problem of learning a performant policy from a fixed batch of data collected by following some behavior policy. Model-based approaches are particularly appealing in the offline setting since they can extract more learning signals from the logged dataset by learning a model of the environment. However, the performance of existing model-based approaches falls short of model-free counterparts, due to the compounding of estimation errors in the learned model. Driven by this observation, we argue that it is critical for a model-based method to understand when to trust the model and when to rely on model-free estimates, and how to act conservatively w.r.t. both. To this end, we derive an elegant and simple methodology called conservative Bayesian model-based value expansion for offline policy optimization (CBOP), that trades off model-free and model-based estimates during the policy evaluation step according to their epistemic uncertainties, and facilitates conservatism by taking a lower bound on the Bayesian posterior value estimate. On the standard D4RL continuous control tasks, we find that our method significantly outperforms previous model-based approaches: e.g., MOPO by %, MOReL by % and COMBO by %. Further, CBOP achieves state-of-the-art performance on out of benchmark datasets while doing on par on the remaining datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Model-Bellman Inconsistency for Model-based Offline Reinforcement LearningYihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin 等ICML 2023 · 被引用 61 次
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 被引用 19 次
- Trust the Model Where It Trusts Itself - Model-Based Actor-Critic with Uncertainty-Aware Rollout AdaptionBernd Frauenknecht, Artur Eisele, Devdutt Subhasish, Friedrich Solowjow 等ICML 2024 · 被引用 14 次
- Scalable Offline Model-Based RL with Action ChunksKwanyoung Park, Seohong Park, Youngwoon Lee, Sergey LevineICLR 2026 · 被引用 12 次
- Bayes Adaptive Monte Carlo Tree Search for Offline Model-based Reinforcement LearningJiayu Chen, Le Xu, Wen-Tse Chen, Jeff SchneiderICLR 2026 · 被引用 10 次
它引用的顶会 Paper15
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
相关 Paper
- OCEAN-MBRL: Offline Conservative Exploration for Model-Based Offline Reinforcement LearningFan Wu, Rui Zhang, Qi Yi, Yunkai Gao 等AAAI 2024 · 被引用 4 次
- Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian LensJihwan Jeong, Xiaoyu Wang, Jingmin Wang, Scott Sanner 等ICML 2025
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran 等NeurIPS 2021 · 被引用 549 次
- VIPO: Value Function Inconsistency Penalized Offline Reinforcement LearningXuyang Chen, Keyu Yan, Guojian Wang, Lin ZhaoICML 2026 · 被引用 3 次
- Model-based Offline Reinforcement Learning with Count-based ConservatismByeongchan Kim, Min-hwan OhICML 2023 · 被引用 19 次
