Learning to Model the World With Language
Jessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner, Pieter Abbeel, Dan Klein, Anca D. Dragan
Abstract
To interact with humans and act in the world, agents need to understand the range of language that people use and relate it to the visual world. While current agents can learn to execute simple language instructions, we aim to build agents that leverage diverse language -- language like"this button turns on the TV"or"I put the bowls away"-- that conveys general knowledge, describes the state of the world, provides interactive feedback, and more. Our key idea is that agents should interpret such diverse language as a signal that helps them predict the future: what they will observe, how the world will behave, and which situations will be rewarded. This perspective unifies language understanding with future prediction as a powerful self-supervised learning objective. We instantiate this in Dynalang, an agent that learns a multimodal world model to predict future text and image representations, and learns to act from imagined model rollouts. While current methods that learn language-conditioned policies degrade in performance with more diverse types of language, we show that Dynalang learns to leverage environment descriptions, game rules, and instructions to excel on tasks ranging from game-playing to navigating photorealistic home scans. Finally, we show that our method enables additional capabilities due to learning a generative model: Dynalang can be pretrained on text-only data, enabling learning from offline datasets, and generate language grounded in an environment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta et al.NeurIPS 2024 · 403 citations
- Generating Code World Models with Large Language Models Guided by Monte Carlo Tree SearchNicola Dainese, Matteo Merler, Minttu Alakuijala, Pekka MarttinenNeurIPS 2024 · 49 citations
- GenRL: Multimodal-foundation world models for generalization in embodied agentsPietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Aaron C. Courville et al.NeurIPS 2024 · 37 citations
- LLM-Empowered State Representation for Reinforcement LearningBoyuan Wang, Yun Qu, Yuhang Jiang, Jianzhun Shao et al.ICML 2024 · 34 citations
- SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied ManipulationJunjie Zhang, Chenjia Bai, Haoran He, Zhigang Wang et al.ICML 2024 · 31 citations
Builds on30
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
Related papers
- ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World ModelsYu-Wei Zhan, Xin Wang, Pengzhe Mao, Tongtong Feng et al.CVPR 2026
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' RuleShuhei Kurita, Kyunghyun ChoICLR 2021 · 29 citations
- RLZero: Direct Policy Inference from Language Without In-Domain SupervisionHarshit Sikchi, Siddhant Agarwal, Pranaya Jajoo, Samyak Parajuli et al.NeurIPS 2025 · 8 citations
- Bridging Environments and Language with Rendering Functions and Vision-Language ModelsThéo Cachet, Christopher R. Dance, Olivier SigaudICML 2024 · 1 citation
- Goal-Aware Prediction: Learning to Model What MattersSuraj Nair, Silvio Savarese, Chelsea FinnICML 2020 · 71 citations
