Model-Based Imaginative Planning for Embodied Agents
Junru Song, Hengzhe Jin, Yucong Huang, Tingsong Jiang, Weien Zhou, Feifei Wang, Yang Yang, Ying Wen, Wen Yao
Abstract
Reasoning and planning critically rely on a predictive dynamics model. In symbolic domains such as mathematics and code, large language models (LLMs) internalize transition rules during pretraining, allowing reinforcement learning or test-time scaling to effectively elicit and generalize their reasoning ability. Embodied decision making is fundamentally different: agents must reason from sparse visual evidence under partial observability, while coping with environment-specific dynamics and affordances not captured by language priors. Here we propose IMPLEMENT, a model-based reasoning framework that enables frozen LLMs to perform imaginative planning. A lightweight world model converts raw pixels into objectcentric symbolic states amenable to languagebased reasoning, and predicts their evolution under hypothetical actions. To address epistemic uncertainty stemming from partial observability, we perform Monte Carlo state prediction via temperature sampling, enabling decision evaluation over multiple plausible futures. To support adaptation to unseen environments, we integrate Meta In-Context Learning, conditioning the world model on interaction history to continually refine its predictions. At inference time, the LLM and world model form a tight co-reasoning loop: the LLM proposes candidate actions, the world model simulates future trajectories, and the LLM refines its decisions, effectively inducing an online policy iteration scheme. Extensive experiments in ALFWorld demonstrate consistent advantages over finetuning-based and strong test-time scaling approaches, validating IMPLEMENT as an effective framework for grounding language agents in visual embodied environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be4bc447-f061-472d-985f-69312ec9b676Builds on14
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
Related papers
- World Model Implanting for Test-time Adaptation of Embodied AgentsMinjong Yoo, Jinwoo Jang, Sihyung Yoon, Honguk WooICML 2025
- PRISM: Perception Reasoning Interleaved for Sequential Decision Making.Mohamed Salim AISSI, Salim Aissi, Clément Romac, Laure Soulier et al.ICML 2026
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language ModelsHaowen Liu, Shaoxiong Yao, Haonan Chen, Jiawei Gao et al.CVPR 2026 · 8 citations
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied ReasoningWonje Choi, Jooyoung Kim, Honguk WooNeurIPS 2025 · 4 citations
- From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous SystemsJianliang He, Siyu Chen, Fengzhuo Zhang, Zhuoran YangICML 2024 · 12 citations
