Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task Agents
Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian Ma, Yitao Liang
Abstract
We investigate the challenge of task planning for multi-task embodied agents in open-world environments. 2 Two main difficulties are identified: 1) executing plans in an open-world environment (e.g., Minecraft) necessitates accurate and multi-step reasoning due to the long-term nature of tasks, and 2) as vanilla planners do not consider how easy the current agent can achieve a given sub-task when ordering parallel sub-goals within a complicated plan, the resulting plan could be inefficient or even infeasible. To this end, we propose "Describe, Explain, Plan and Select" (DEPS), an interactive planning approach based on Large Language Models (LLMs). DEPS facilitates better error correction on initial LLM-generated plan by integrating description of the plan execution process and providing selfexplanation of feedback when encountering failures during the extended planning phases. Furthermore, it includes a goal selector, which is a trainable module that ranks parallel candidate sub-goals based on the estimated steps of completion, consequently refining the initial plan. Our experiments mark the milestone of the first zero-shot multi-task agent that can robustly accomplish 70+ Minecraft tasks and nearly double the overall performances. Further testing reveals our method's general effectiveness in popularly adopted non-open-ended domains as well (i.e., ALFWorld and tabletop manipulation). The ablation and exploratory studies detail how our design beats the counterparts and provide a promising update on the ObtainDiamond grand challenge with our approach. The code is released at https://github.com/CraftJarvis/MC-Planner . * Corresponding Author. 2 We borrow the term "open world" from the game community. It highlights that the agent can navigate inside a diverse environment and accomplish open-ended tasks freely. 37th Conference on Neural Information Processing Systems (NeurIPS 2023). (Re-)Planner LLM * Controller Goal-conditioned Policy Selector HPM Descriptor VLM Explainer LLM * Instruction plan ๐ ! goal ๐ ! feedback action obs description ๐ ! explain Environment obs Task instruction: Obtain a diamond in Minecraft survival mode step-by-step? Candidate goals: Selected Goal ๐ ๐ : ร 4 The agent locates in the birch forest, which only has birch wood. Description ๐ ๐ : I succeed on goal 1-5. I fail on goal 6, mining 3 with . Now my inventory has 5 planks, โฆ Initial Plan ๐ท ๐ :
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 384e1799-b03f-4c8a-a07c-232bd94e04e3Cited by top-tier papers31
- DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based ReasoningSiyuan Guo, Cheng Deng, Ying Wen, Hechang Chen et al.ICML 2024 ยท 107 citations
- SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D PriorsChenyang Ma, Kai Lu, Ta Ying Cheng, Niki Trigoni et al.NeurIPS 2024 ยท 82 citations
- Richelieu: Self-Evolving LLM-Based Agents for AI DiplomacyZhenyu Guan, Xiangyu Kong, Fangwei Zhong, Yizhou WangNeurIPS 2024 ยท 48 citations
- ARTiST: Automated Text Simplification for Task Guidance in Augmented RealityGuande Wu, Jing Qian, Sonia Castelo Quispe, Shaoyu Chen et al.CHI 2024 ยท 20 citations
- Planning in the Dark: LLM-Symbolic Planning Pipeline Without ExpertsSukai Huang, Nir Lipovetzky, Trevor CohnAAAI 2025 ยท 19 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 ยท 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 ยท 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 ยท 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 ยท 22,562 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 ยท 6,707 citations
Related papers
- MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active PerceptionYiran Qin, Enshen Zhou, Qichang Liu, Zhenfei Yin et al.CVPR 2024 ยท 11 citations
- Do Embodied Agents Dream of Pixelated Sheep: Embodied Decision Making using Language Guided World ModellingKolby Nottingham, Prithviraj Ammanabrolu, Alane Suhr, Yejin Choi et al.ICML 2023 ยท 110 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 ยท 685 citations
- Experience-based Knowledge Correction for Robust Planning in MinecraftSeungjoon Lee, Suhwan Kim, Minhyeon Oh, Youngsik Yoon et al.ICLR 2026 ยท 1 citation
- Interactive and Expressive Code-Augmented Planning with Large Language ModelsAnthony Zhe Liu, Xinhe Wang, Jacob Sansom, Yao Fu et al.ACL 2025 ยท 4 citations
