Infer Human's Intentions Before Following Natural Language Instructions
Yanming Wan, Yue Wu, Yiping Wang, Jiayuan Mao, Natasha Jaques
Abstract
For AI agents to be helpful to humans, they should be able to follow natural language instructions to complete everyday cooperative tasks in human environments. However, real human instructions inherently possess ambiguity, because the human speakers assume sufficient prior knowledge about their hidden goals and intentions. Standard language grounding and planning methods fail to address such ambiguities because they do not model human internal goals as additional partially observable factors in the environment. We propose a new framework, Follow Instructions with Social and Embodied Reasoning (FISER) * , aiming for better natural language instruction following in collaborative embodied tasks. Our framework makes explicit inferences about human goals and intentions as intermediate reasoning steps. We implement a set of Transformer-based models and evaluate them over a challenging benchmark, HandMeThat. We empirically demonstrate that using social reasoning to explicitly infer human intentions before making action plans surpasses purely end-toend approaches. We also compare our implementation with strong baselines, including Chain of Thought prompting on the largest available pre-trained language models, and find that FISER provides better performance on the embodied social reasoning tasks under investigation, reaching the state-of-theart on HandMeThat.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f611ac3d-059a-4036-b622-dcd227a5673fCited by top-tier papers2
- COOPERA: Continual Open-Ended Human-Robot AssistanceChenyang Ma, Kai Lu, Ruta Desai, Xavier Puig et al.NeurIPS 2025 · 9 citations
- Modeling Others' Minds as CodeKunal Jha, Aydan Yuenan Huang, Eric Ye, Natasha Jaques et al.ICLR 2026 · 6 citations
Builds on6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- TEACh: Task-Driven Embodied Agents That ChatAishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange et al.AAAI 2022 · 251 citations
- Watch-And-Help: A Challenge for Social Perception and Human-AI CollaborationXavier Puig, Tianmin Shu, Shuang Li, Zilin Wang et al.ICLR 2021 · 170 citations
- Neural Symbolic Reader: Scalable Integration of Distributed and Symbolic Representations for Reading ComprehensionXinyun Chen, Chen Liang, Adams Wei Yu, Denny Zhou et al.ICLR 2020 · 109 citations
- Enhancing Human-AI Collaboration Through Logic-Guided ReasoningChengzhi Cao, Yinghao Fu, Sheng Xu, Ruimao Zhang et al.ICLR 2024 · 7 citations
Related papers
- ThinkBot: Embodied Instruction Following with Thought Chain ReasoningGuanxing Lu, Ziwei Wang, Changliu Liu, Jiwen Lu et al.ICLR 2025
- Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few ExamplesTaewoong Kim, Byeonghwi Kim, Jonghyun ChoiAAAI 2025 · 8 citations
- DANLI: Deliberative Agent for Following Natural Language InstructionsYichi Zhang, Jianing Yang, Jiayi Pan, Shane Storks et al.EMNLP 2022 · 4 citations
- Language Instructed Reinforcement Learning for Human-AI CoordinationHengyuan Hu, Dorsa SadighICML 2023 · 90 citations
- Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied DialogueAishwarya Padmakumar, Mert Inan, Spandana Gella, Patrick Lange et al.EMNLP 2023 · 1 citation
