On Effective Scheduling of Model-based Reinforcement Learning
Hang Lai, Jian Shen, Weinan Zhang, Yimin Huang, Xing Zhang, Ruiming Tang, Yong Yu, Zhenguo Li
Abstract
Model-based reinforcement learning has attracted wide attention due to its superior sample efficiency. Despite its impressive success so far, it is still unclear how to appropriately schedule the important hyperparameters to achieve adequate performance, such as the real data ratio for policy optimization in Dyna-style model-based algorithms. In this paper, we first theoretically analyze the role of real data in policy training, which suggests that gradually increasing the ratio of real data yields better performance. Inspired by the analysis, we propose a framework named AutoMBPO to automatically schedule the real data ratio as well as other hyperparameters in training model-based policy optimization (MBPO) algorithm, a representative running case of model-based methods. On several continuous control tasks, the MBPO instance trained with hyperparameters scheduled by AutoMBPO can significantly surpass the original one, and the real data ratio schedule found by AutoMBPO shows consistency with our theoretical analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- When to Update Your Model: Constrained Model-based Reinforcement LearningTianying Ji, Yu Luo, Fuchun Sun, Mingxuan Jing et al.NeurIPS 2022 · 27 citations
- Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-CriticTianying Ji, Yu Luo, Fuchun Sun, Xianyuan Zhan et al.ICML 2024 · 23 citations
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 20 citations
- Learning World Models for Unconstrained Goal NavigationYuanlin Duan, Wensen Mao, He ZhuNeurIPS 2024 · 11 citations
- Is Model Ensemble Necessary? Model-based RL via a Single Model with Lipschitz Regularized Value FunctionRuijie Zheng, Xiyao Wang, Huazhe Xu, Furong HuangICLR 2023
Builds on6
- A Game Theoretic Framework for Model Based Reinforcement LearningAravind Rajeswaran, Igor Mordatch, Vikash KumarICML 2020 · 137 citations
- Model-Augmented Actor-Critic: Backpropagating through PathsIgnasi Clavera, Yao Fu, Pieter AbbeelICLR 2020 · 96 citations
- Bidirectional Model-based Policy OptimizationHang Lai, Jian Shen, Weinan Zhang, Yong YuICML 2020 · 66 citations
- Trust the Model When It Is Confident: Masked Model-based Actor-CriticFeiyang Pan, Jia He, Dandan Tu, Qing HeNeurIPS 2020 · 65 citations
- Model-based Policy Optimization with Unsupervised Model AdaptationJian Shen, Han Zhao, Weinan Zhang, Yong YuNeurIPS 2020 · 33 citations
Related papers
- Using Probabilistic Model Rollouts to Boost the Sample Efficiency of Reinforcement Learning for Automated Analog Circuit SizingMohsen Ahmadzadeh, Georges G. E. GielenDAC 2024 · 5 citations
- On Rollouts in Model-Based Reinforcement LearningBernd Frauenknecht, Devdutt Subhasish, Friedrich Solowjow, Sebastian TrimpeICLR 2025 · 1 citation
- Making Better Decision by Directly Planning in Continuous ControlJinhua Zhu, Yue Wang, Lijun Wu, Tao Qin et al.ICLR 2023
- Model-Based Offline PlanningArthur Argenson, Gabriel Dulac-ArnoldICLR 2021 · 20 citations
- Sample-Efficient Automated Deep Reinforcement LearningJörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank HutterICLR 2021 · 49 citations
