The Benefits of Model-Based Generalization in Reinforcement Learning
Kenny John Young, Aditya A. Ramesh, Louis Kirsch, Jürgen Schmidhuber
Abstract
Model-Based Reinforcement Learning (RL) is widely believed to have the potential to improve sample efficiency by allowing an agent to synthesize large amounts of imagined experience. Experience Replay (ER) can be considered a simple kind of model, which has proved effective at improving the stability and efficiency of deep RL. In principle, a learned parametric model could improve on ER by generalizing from real experience to augment the dataset with additional plausible experience. However, given that learned value functions can also generalize, it is not immediately obvious why model generalization should be better. Here, we provide theoretical and empirical insight into when, and how, we can expect data generated by a learned model to be useful. First, we provide a simple theorem motivating how learning a model as an intermediate step can narrow down the set of possible value functions more than learning a value function directly from data using the Bellman equation. Second, we provide an illustrative example showing empirically how a similar effect occurs in a more concrete setting with neural network function approximation. Finally, we provide extensive experiments showing the benefit of model-based learning for online RL in environments with combinatorial complexity, but factored structure that allows a learned model to generalize. In these experiments, we take care to control for other factors in order to isolate, insofar as possible, the benefit of using experience generated by a learned model relative to ER alone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8476fa82-0c0a-4c62-baf7-00af1c7b2941Cited by top-tier papers7
- Closing the Gap between TD Learning and Supervised Learning - A Generalisation Point of ViewRaj Ghugare, Matthieu Geist, Glen Berseth, Benjamin EysenbachICLR 2024 · 28 citations
- Distributional Successor Features Enable Zero-Shot Policy OptimizationChuning Zhu, Xinqi Wang, Tyler Han, Simon S. Du et al.NeurIPS 2024 · 11 citations
- Learning World Models for Unconstrained Goal NavigationYuanlin Duan, Wensen Mao, He ZhuNeurIPS 2024 · 11 citations
- Policy Rehearsing: Training Generalizable Policies for Reinforcement LearningChengxing Jia, Chenxiao Gao, Hao Yin, Fuxiang Zhang et al.ICLR 2024 · 6 citations
- QORA: Zero-Shot Transfer via Interpretable Object-Relational Model LearningGabriel Stella, Dmitri LoguinovICML 2024 · 1 citation
Builds on10
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel et al.NeurIPS 2021 · 345 citations
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 120 citations
Related papers
- Prioritized Generative ReplayRenhao Wang, Kevin Frans, Pieter Abbeel, Sergey Levine et al.ICLR 2025
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez et al.ICLR 2021 · 77 citations
- Theoretically Principled Deep RL Acceleration via Nearest Neighbor Function ApproximationJunhong Shen, Lin F. YangAAAI 2021 · 29 citations
- Latent Variable Representation for Reinforcement LearningTongzheng Ren, Chenjun Xiao, Tianjun Zhang, Na Li et al.ICLR 2023 · 1 citation
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
