On-Policy Model Errors in Reinforcement Learning
Lukas P. Fröhlich, Maksym Lefarov, Melanie N. Zeilinger, Felix Berkenkamp
摘要
Model-free reinforcement learning algorithms can compute policy gradients given sampled environment transitions, but require large amounts of data. In contrast, model-based methods can use the learned model to generate new data, but model errors and bias can render learning unstable or suboptimal. In this paper, we present a novel method that combines real-world data and a learned model in order to get the best of both worlds. The core idea is to exploit the real-world data for on-policy predictions and use the learned model only to generalize to different actions. Specifically, we use the data as time-dependent on-policy correction terms on top of a learned model, to retain the ability to generate data without accumulating errors over long prediction horizons. We motivate this method theoretically and show that it counteracts an error term for model-based policy improvement. Experiments on MuJoCo- and PyBullet-benchmarks show that our method can drastically improve existing model-based approaches without introducing additional tuning parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 被引用 20 次
- How to Fine-tune the Model: Unified Model Shift and Model Bias Policy OptimizationHai Zhang, Hang Yu, Junqiao Zhao, Di Zhang 等NeurIPS 2023 · 被引用 16 次
- Is Model Ensemble Necessary? Model-based RL via a Single Model with Lipschitz Regularized Value FunctionRuijie Zheng, Xiyao Wang, Huazhe Xu, Furong HuangICLR 2023
它引用的顶会 Paper4
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 被引用 120 次
- Model-Augmented Actor-Critic: Backpropagating through PathsIgnasi Clavera, Yao Fu, Pieter AbbeelICLR 2020 · 被引用 96 次
相关 Paper
- Offline Data Enhanced On-Policy Policy Gradient with Provable GuaranteesYifei Zhou, Ayush Sekhari, Yuda Song, Wen SunICLR 2024 · 被引用 11 次
- Evaluating Model-Based Planning and Planner Amortization for Continuous ControlArunkumar Byravan, Leonard Hasenclever, Piotr Trochim, Mehdi Mirza 等ICLR 2022 · 被引用 18 次
- Dream-MPC: Gradient-Based Model Predictive Control with Latent ImaginationJonathan Spieler, Sven BehnkeICML 2026
- PWM: Policy Learning with Multi-Task World ModelsIgnat Georgiev, Varun Giridhar, Nicklas Hansen, Animesh GargICLR 2025
- Bidirectional Model-based Policy OptimizationHang Lai, Jian Shen, Weinan Zhang, Yong YuICML 2020 · 被引用 66 次
