ThinkBot: Embodied Instruction Following with Thought Chain Reasoning
Guanxing Lu, Ziwei Wang, Changliu Liu, Jiwen Lu, Yansong Tang
摘要
Embodied Instruction Following (EIF) requires agents to complete human instruction by interacting objects in complicated surrounding environments. Conventional methods directly consider the sparse human instruction to generate action plans for agents, which usually fail to achieve human goals because of the instruction incoherence in action descriptions. On the contrary, we propose ThinkBot that reasons the thought chain in human instruction to recover the missing action descriptions, so that the agent can successfully complete human goals by following the coherent instruction. Specifically, we first design an instruction completer based on large language models to recover the missing actions with interacted objects between consecutive human instruction, where the perceived surrounding environments and the completed sub-goals are considered for instruction completion. Based on the partially observed scene semantic maps, we present an object localizer to infer the position of interacted objects for agents to achieve complex human goals. Extensive experiments in the simulated environment show that our ThinkBot outperforms the stateof-the-art EIF methods by a sizable margin in both success rate and execution efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningSiyin Wang, Zhaoye Fei, Qinyuan Cheng, Shiduo Zhang 等ACL 2025 · 被引用 16 次
- Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction FollowingYueen Ma, Dafeng Chi, Shiguang Wu, Yuecheng Liu 等EMNLP 2025 · 被引用 9 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
相关 Paper
- OPEx: A Component-Wise Analysis of LLM-Centric Agents in Embodied Instruction FollowingHaochen Shi, Zhiyuan Sun, Xingdi Yuan, Marc-Alexandre Côté 等ACL 2024
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu 等NeurIPS 2025 · 被引用 18 次
- Infer Human's Intentions Before Following Natural Language InstructionsYanming Wan, Yue Wu, Yiping Wang, Jiayuan Mao 等AAAI 2025 · 被引用 10 次
- Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied DialogueAishwarya Padmakumar, Mert Inan, Spandana Gella, Patrick Lange 等EMNLP 2023 · 被引用 1 次
- Human-Object Interaction from Human-level InstructionsZhen Wu, Jiaman Li, Pei Xu, C. Karen LiuICCV 2025 · 被引用 3 次
