ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, Matthew J. Hausknecht
Abstract
Given a simple request (e.g., Put a washed apple in the kitchen fridge), humans can reason in purely abstract terms by imagining action sequences and scoring their likelihood of success, prototypicality, and efficiency, all without moving a muscle. Once we see the kitchen in question, we can update our abstract plans to fit the scene. Embodied agents require the same abilities, but existing work does not yet provide the infrastructure necessary for both reasoning abstractly and executing concretely. We address this limitation by introducing ALFWorld, a simulator that enables agents to learn abstract, text-based policies in TextWorld (Cote et al., 2018) and then execute goals from the ALFRED benchmark (Shridhar et al., 2020) in a rich visual environment. ALFWorld enables the creation of a new BUTLER agent whose abstract knowledge, learned in TextWorld, corresponds directly to concrete, visually grounded actions. In turn, as we demonstrate empirically, this fosters better agent generalization than training only in the visually grounded environment. BUTLER's simple, modular design factors the problem to allow researchers to focus on models for improving every piece of the pipeline (language understanding, planning, navigation, visual scene understanding, and so forth).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1916a0c5-304d-4662-aee9-7bc796e5e6d6Cited by top-tier papers245
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
Builds on8
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu et al.NeurIPS 2020 · 251 citations
- Interactive Fiction Games: A Colossal AdventureMatthew J. Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, Xingdi YuanAAAI 2020 · 242 citations
- Graph Constrained Reinforcement Learning for Natural Language Action SpacesPrithviraj Ammanabrolu, Matthew J. HausknechtICLR 2020 · 138 citations
- Learning Dynamic Belief Graphs to Generalize on Text-Based GamesAshutosh Adhikari, Xingdi Yuan, Marc-Alexandre Côté, Mikulas Zelinka et al.NeurIPS 2020 · 91 citations
Related papers
- Model-Based Imaginative Planning for Embodied AgentsJunru Song, Hengzhe Jin, Yucong Huang, Tingsong Jiang et al.ACL 2026
- Causal Abstraction Learning for Multi-Modal Grounded PlanningXinshu Li, Shiyi Yang, Ziqi Xu, Feng Xia et al.KDD 2026
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk et al.CVPR 2020
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao et al.CVPR 2024 · 23 citations
- Learning Grounded Action Abstractions from LanguageLionel Wong, Jiayuan Mao, Pratyusha Sharma, Zachary S. Siegel et al.ICLR 2024 · 7 citations
