How to Fine-tune the Model: Unified Model Shift and Model Bias Policy Optimization
Hai Zhang, Hang Yu, Junqiao Zhao, Di Zhang, Xiao Zhang, Hongtu Zhou, Chang Huang, Chen Ye
Abstract
Designing and deriving effective model-based reinforcement learning (MBRL) algorithms with a performance improvement guarantee is challenging, mainly attributed to the high coupling between model learning and policy optimization. Many prior methods that rely on return discrepancy to guide model learning ignore the impacts of model shift, which can lead to performance deterioration due to excessive model updates. Other methods use performance difference bound to explicitly consider model shift. However, these methods rely on a fixed threshold to constrain model shift, resulting in a heavy dependence on the threshold and a lack of adaptability during the training process. In this paper, we theoretically derive an optimization objective that can unify model shift and model bias and then formulate a fine-tuning process. This process adaptively adjusts the model updates to get a performance improvement guarantee while avoiding model overfitting. Based on these, we develop a straightforward algorithm USB-PO 2 (Unified model Shift and model Bias Policy Optimization). Empirical results show that USB-PO achieves state-of-the-art performance on several challenging benchmark tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5640322-8355-4a23-8788-d2f79e44a144Cited by top-tier papers5
- Pixel to Gaussian: Ultra-Fast Continuous Super-Resolution with 2D Gaussian ModelingLong Peng, Anran Wu, Wenbo Li, Peizhe Xia et al.ICLR 2026 · 57 citations
- Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement LearningLanqing Li, Hai Zhang, Xinyu Zhang, Shatong Zhu et al.NeurIPS 2024 · 24 citations
- Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-CriticTianying Ji, Yu Luo, Fuchun Sun, Xianyuan Zhan et al.ICML 2024 · 23 citations
- Scrutinize What We Ignore: Reining In Task Representation Shift Of Context-Based Offline Meta Reinforcement LearningHai Zhang, Boyuan Zheng, Tianying Ji, Jinhang Liu et al.ICLR 2025
- Highly Efficient Self-Adaptive Reward Shaping for Reinforcement LearningHaozhe Ma, Zhengding Luo, Thanh Vinh Vo, Kuankuan Sima et al.ICLR 2025
Builds on17
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain et al.NeurIPS 2021 · 149 citations
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 120 citations
- Bidirectional Model-based Policy OptimizationHang Lai, Jian Shen, Weinan Zhang, Yong YuICML 2020 · 66 citations
Related papers
- When to Update Your Model: Constrained Model-based Reinforcement LearningTianying Ji, Yu Luo, Fuchun Sun, Mingxuan Jing et al.NeurIPS 2022 · 27 citations
- The Virtues of Laziness in Model-based RL: A Unified Objective and AlgorithmsAnirudh Vemula, Yuda Song, Aarti Singh, Drew Bagnell et al.ICML 2023 · 15 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
- Deep Model-Based Reinforcement Learning via Estimated Uncertainty and Conservative Policy OptimizationQi Zhou, Houqiang Li, Jie WangAAAI 2020 · 17 citations
- Model-based Policy Optimization with Unsupervised Model AdaptationJian Shen, Han Zhao, Weinan Zhang, Yong YuNeurIPS 2020 · 33 citations
