Procedural generalization by planning with self-supervised world models
Ankesh Anand, Jacob C. Walker, Yazhe Li, Eszter Vértes, Julian Schrittwieser, Sherjil Ozair, Theophane Weber, Jessica B. Hamrick
Abstract
One of the key promises of model-based reinforcement learning is the ability to generalize using an internal model of the world to make predictions in novel environments and tasks. However, the generalization ability of model-based agents is not well understood because existing work has focused on model-free agents when benchmarking generalization. Here, we explicitly measure the generalization ability of model-based agents in comparison to their model-free counterparts. We focus our analysis on MuZero [60], a powerful model-based agent, and evaluate its performance on both procedural and task generalization. We identify three factors of procedural generalization-planning, self-supervised representation learning, and procedural data diversity-and show that by combining these techniques, we achieve state-of-the art generalization performance and data efficiency on Procgen [9] . However, we find that these factors do not always provide the same benefits for the task generalization benchmarks in Meta-World [74] , indicating that transfer remains a challenge and may require different approaches than procedural generalization. Overall, we suggest that building generalizable agents requires moving beyond the single-task, model-free paradigm and towards self-supervised model-based agents that are trained in rich, procedural, multi-task environments. * Work done while visiting from Mila, University of Montreal. † Joint first authors. ‡ Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- On the Importance of Exploration for Generalization in Reinforcement LearningYiding Jiang, J. Zico Kolter, Roberta RaileanuNeurIPS 2023 · 48 citations
- On the Effectiveness of Fine-tuning Versus Meta-reinforcement LearningMandi Zhao, Pieter Abbeel, Stephen JamesNeurIPS 2022 · 43 citations
- Grounding Multimodal Large Language Models in ActionsAndrew Szot, Bogdan Mazoure, Harsh Agrawal, R. Devon Hjelm et al.NeurIPS 2024 · 43 citations
- What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?Rui Yang, Lin Yong, Xiaoteng Ma, Hao Hu et al.ICML 2023 · 35 citations
- Efficient World Models with Context-Aware TokenizationVincent Micheli, Eloi Alonso, François FleuretICML 2024 · 25 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
Related papers
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez et al.ICLR 2021 · 77 citations
- Investigating the Role of Model-Based Learning in Exploration and TransferJacob C. Walker, Eszter Vértes, Yazhe Li, Gabriel Dulac-Arnold et al.ICML 2023 · 8 citations
- Efficient Multi-agent Reinforcement Learning by PlanningQihan Liu, Jianing Ye, Xiaoteng Ma, Jun Yang et al.ICLR 2024 · 18 citations
- Planning in Stochastic Environments with a Learned ModelIoannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K. Hubert et al.ICLR 2022 · 79 citations
- Meta-Reinforcement Learning Based on Self-Supervised Task Representation LearningMingyang Wang, Zhenshan Bing, Xiangtong Yao, Shuai Wang et al.AAAI 2023 · 22 citations
