ThinkBot: Embodied Instruction Following with Thought Chain Reasoning
Guanxing Lu, Ziwei Wang, Changliu Liu, Jiwen Lu, Yansong Tang
Abstract
Embodied Instruction Following (EIF) requires agents to complete human instruction by interacting objects in complicated surrounding environments. Conventional methods directly consider the sparse human instruction to generate action plans for agents, which usually fail to achieve human goals because of the instruction incoherence in action descriptions. On the contrary, we propose ThinkBot that reasons the thought chain in human instruction to recover the missing action descriptions, so that the agent can successfully complete human goals by following the coherent instruction. Specifically, we first design an instruction completer based on large language models to recover the missing actions with interacted objects between consecutive human instruction, where the perceived surrounding environments and the completed sub-goals are considered for instruction completion. Based on the partially observed scene semantic maps, we present an object localizer to infer the position of interacted objects for agents to achieve complex human goals. Extensive experiments in the simulated environment show that our ThinkBot outperforms the stateof-the-art EIF methods by a sizable margin in both success rate and execution efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ce7cc92-45cf-43b8-8493-51d13a4f5f7cCited by top-tier papers2
- World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningSiyin Wang, Zhaoye Fei, Qinyuan Cheng, Shiduo Zhang et al.ACL 2025 · 16 citations
- Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction FollowingYueen Ma, Dafeng Chi, Shiguang Wu, Yuecheng Liu et al.EMNLP 2025 · 9 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
Related papers
- OPEx: A Component-Wise Analysis of LLM-Centric Agents in Embodied Instruction FollowingHaochen Shi, Zhiyuan Sun, Xingdi Yuan, Marc-Alexandre Côté et al.ACL 2024
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu et al.NeurIPS 2025 · 18 citations
- Infer Human's Intentions Before Following Natural Language InstructionsYanming Wan, Yue Wu, Yiping Wang, Jiayuan Mao et al.AAAI 2025 · 10 citations
- Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied DialogueAishwarya Padmakumar, Mert Inan, Spandana Gella, Patrick Lange et al.EMNLP 2023 · 1 citation
- Human-Object Interaction from Human-level InstructionsZhen Wu, Jiaman Li, Pei Xu, C. Karen LiuICCV 2025 · 3 citations
