Model-based Offline RL via Robust Value-Aware Model Learning with Implicitly Differentiable Adaptive Weighting
Zhongjian Qiao, Jiafei Lyu, Boxiang Lyu, Yao Shu, Siyang Gao, Shuang Qiu
Abstract
Model-based offline reinforcement learning (RL) aims to enhance offline RL with a dynamics model that facilitates policy exploration. However, model exploitation could occur due to inevitable model errors, degrading algorithm performance. Adversarial model learning offers a theoretical framework to mitigate model exploitation by solving a maximin formulation. Within such a paradigm, RAMBO has emerged as a representative and most popular method that provides a practical implementation with model gradient. However, we empirically reveal that severe Q-value underestimation and gradient explosion can occur in RAMBO with only slight hyperparameter tuning, suggesting that it tends to be overly conservative and suffers from unstable model updates. To address these issues, we propose RObust value-aware Model learning with Implicitly differentiable adaptive weighting (ROMI). Instead of updating the dynamics model with model gradient, ROMI introduces a novel robust value-aware model learning approach. This approach requires the dynamics model to predict future states with values close to the minimum Q-value within a scale-adjustable state uncertainty set, enabling controllable conservatism and stable model updates. To further improve out-of-distribution (OOD) generalization during multi-step rollouts, we propose implicitly differentiable adaptive weighting, a bi-level optimization scheme that adaptively achieves dynamics- and value-aware model learning. Empirical results on D4RL and NeoRL datasets show that ROMI significantly outperforms RAMBO and achieves competitive or superior performance compared to other state-of-the-art methods on datasets where RAMBO typically underperforms. Code is available at https://github.com/zq2r/ROMI.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f0f226f-97ae-4055-a068-fe11bf2f7390Cited by top-tier papers4
- Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics ShiftsZhongjian Qiao, Rui Yang, Jiafei Lyu, Xiu Li et al.ICLR 2026 · 7 citations
- Transport or Discard: Robust Unbalanced Optimal Transport for Cross-Domain Policy AdaptationWenyu Chen, Yujia Zhang, Wei Guo, Linli Ma et al.ICML 2026
- : Decision-Targeted Digital TwinsHarry Amad, Mihaela van der SchaarICML 2026
- Unifying Value Alignment and Assignment in Cross-Domain Offline Reinforcement Learning with Heterogeneous DatasetsZhongjian Qiao, Jiafei Lyu, Chenjia Bai, Peisong Wang et al.ICML 2026
Builds on32
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
Related papers
- OCEAN-MBRL: Offline Conservative Exploration for Model-Based Offline Reinforcement LearningFan Wu, Rui Zhang, Qi Yi, Yunkai Gao et al.AAAI 2024 · 4 citations
- RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2022 · 168 citations
- Model-Bellman Inconsistency for Model-based Offline Reinforcement LearningYihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin et al.ICML 2023 · 61 citations
- Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement LearningJiayu Chen, Le Xu, Aravind Venugopal, Jeff SchneiderICML 2026
- Offline Reinforcement Learning with Reverse Model-based ImaginationJianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu et al.NeurIPS 2021 · 74 citations
