When to Update Your Model: Constrained Model-based Reinforcement Learning
Tianying Ji, Yu Luo, Fuchun Sun, Mingxuan Jing, Fengxiang He, Wenbing Huang
摘要
Designing and analyzing model-based RL (MBRL) algorithms with guaranteed monotonic improvement has been challenging, mainly due to the interdependence between policy optimization and model learning. Existing discrepancy bounds generally ignore the impacts of model shifts, and their corresponding algorithms are prone to degrade performance by drastic model updating. In this work, we first propose a novel and general theoretical scheme for a non-decreasing performance guarantee of MBRL. Our follow-up derived bounds reveal the relationship between model shifts and performance improvement. These discoveries encourage us to formulate a constrained lower-bound optimization problem to permit the monotonicity of MBRL. A further example demonstrates that learning models from a dynamically-varying number of explorations benefit the eventual returns. Motivated by these analyses, we design a simple but effective algorithm CMLO 1 (Constrained Model-shift Lower-bound Optimization), by introducing an event-triggered mechanism that flexibly determines when to update the model. Experiments show that CMLO surpasses other state-of-the-art methods and produces a boost when various policy optimization methods are employed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-CriticTianying Ji, Yu Luo, Fuchun Sun, Xianyuan Zhan 等ICML 2024 · 被引用 23 次
- How to Fine-tune the Model: Unified Model Shift and Model Bias Policy OptimizationHai Zhang, Hang Yu, Junqiao Zhao, Di Zhang 等NeurIPS 2023 · 被引用 16 次
- Trust the Model Where It Trusts Itself - Model-Based Actor-Critic with Uncertainty-Aware Rollout AdaptionBernd Frauenknecht, Artur Eisele, Devdutt Subhasish, Friedrich Solowjow 等ICML 2024 · 被引用 14 次
- Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RLYu Luo, Tianying Ji, Fuchun Sun, Jianwei Zhang 等ICML 2024 · 被引用 7 次
- Learning to Play Atari in a World of TokensPranav Agarwal, Sheldon Andrews, Samira Ebrahimi KahouICML 2024 · 被引用 6 次
它引用的顶会 Paper6
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative ModelGen Li, Yuting Wei, Yuejie Chi, Yuantao Gu 等NeurIPS 2020 · 被引用 159 次
- Information Theoretic Regret Bounds for Online Nonlinear ControlSham M. Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi 等NeurIPS 2020 · 被引用 137 次
- Trust the Model When It Is Confident: Masked Model-based Actor-CriticFeiyang Pan, Jia He, Dandan Tu, Qing HeNeurIPS 2020 · 被引用 65 次
- Model-based Reinforcement Learning for Continuous Control with Posterior SamplingYing Fan, Yifei MingICML 2021 · 被引用 25 次
相关 Paper
- Scrutinize What We Ignore: Reining In Task Representation Shift Of Context-Based Offline Meta Reinforcement LearningHai Zhang, Boyuan Zheng, Tianying Ji, Jinhang Liu 等ICLR 2025
- Monotonic Robust Policy Optimization with Model DiscrepancyYuankun Jiang, Chenglin Li, Wenrui Dai, Junni Zou 等ICML 2021 · 被引用 24 次
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 被引用 18 次
- Robust Model Based Reinforcement Learning Using L1 Adaptive ControlMinjun Sung, Sambhu H. Karumanchi, Aditya Gahlawat, Naira HovakimyanICLR 2024 · 被引用 1 次
- Conservative Dual Policy Optimization for Efficient Model-Based Reinforcement LearningShenao ZhangNeurIPS 2022 · 被引用 8 次
