RTFM: Generalising to New Environment Dynamics via Reading
Victor Zhong, Tim Rocktäschel, Edward Grefenstette
Abstract
Obtaining policies that can generalise to new environments in reinforcement learning is challenging. In this work, we demonstrate that language understanding via a reading policy learner is a promising vehicle for generalisation to new environments. We propose a grounded policy learning problem, Read to Fight Monsters (RTFM), in which the agent must jointly reason over a language goal, relevant dynamics described in a document, and environment observations. We procedurally generate environment dynamics and corresponding language descriptions of the dynamics, such that agents must read to understand new environment dynamics instead of memorising any particular information. In addition, we propose txt2π, a model that captures three-way interactions between the goal, document, and observations. On RTFM, txt2π generalises to new environments with dynamics not seen during training via reading. Furthermore, our model outperforms baselines such as FiLM and language-conditioned CNNs on RTFM. Through curriculum learning, txt2π produces policies that excel on complex RTFM tasks requiring several reasoning and coreference steps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu et al.NeurIPS 2020 · 251 citations
- Multi-Task Reinforcement Learning with Context-based RepresentationsShagun Sodhani, Amy Zhang, Joelle PineauICML 2021 · 241 citations
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
- Do Embodied Agents Dream of Pixelated Sheep: Embodied Decision Making using Language Guided World ModellingKolby Nottingham, Prithviraj Ammanabrolu, Alane Suhr, Yejin Choi et al.ICML 2023 · 110 citations
- Learning to Model the World With LanguageJessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner et al.ICML 2024 · 76 citations
Related papers
- Reader: Model-based language-instructed reinforcement learningNicola Dainese, Pekka Marttinen, Alexander IlinEMNLP 2023 · 1 citation
- Conceptual Reinforcement Learning for Language-Conditioned TasksShaohui Peng, Xing Hu, Rui Zhang, Jiaming Guo et al.AAAI 2023 · 12 citations
- Grounding Language to Entities and Dynamics for Generalization in Reinforcement LearningAustin W. Hanjie, Victor Zhong, Karthik NarasimhanICML 2021 · 60 citations
- Learning Knowledge Graph-based World Models of Textual EnvironmentsPrithviraj Ammanabrolu, Mark O. RiedlNeurIPS 2021 · 43 citations
- Improving Policy Learning via Language Dynamics DistillationVictor Zhong, Jesse Mu, Luke Zettlemoyer, Edward Grefenstette et al.NeurIPS 2022 · 16 citations
