Multi-Level Compositional Reasoning for Interactive Instruction Following
Suvaansh Bhambri, Byeonghwi Kim, Jonghyun Choi
Abstract
Robotic agents performing domestic chores by natural language directives are required to master the complex job of navigating environment and interacting with objects in the environments. The tasks given to the agents are often composite thus are challenging as completing them require to reason about multiple subtasks, e.g., bring a cup of coffee. To address the challenge, we propose to divide and conquer it by breaking the task into multiple subgoals and attend to them individually for better navigation and interaction. We call it Multi-level Compositional Reasoning Agent (MCR-Agent). Specifically, we learn a three-level action policy. At the highest level, we infer a sequence of human-interpretable subgoals to be executed based on language instructions by a high-level policy composition controller. At the middle level, we discriminatively control the agent’s navigation by a master policy by alternating between a navigation policy and various independent interaction policies. Finally, at the lowest level, we infer manipulation actions with the corresponding object masks using the appropriate interaction policy. Our approach not only generates human interpretable subgoals but also achieves 2.03% absolute gain to comparable state of the arts in the efficiency metric (PLWSR in unseen set) without using rule-based planning or a semantic spatial memory. The code is available at https://github.com/yonseivnl/mcr-agent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied AgentsByeonghwi Kim, Jinyeon Kim, Yuyeong Kim, Cheolhong Min et al.ICCV 2023 · 46 citations
- Online Continual Learning for Interactive Instruction Following AgentsByeonghwi Kim, Minhyuk Seo, Jonghyun ChoiICLR 2024 · 22 citations
- Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few ExamplesTaewoong Kim, Byeonghwi Kim, Jonghyun ChoiAAAI 2025 · 8 citations
Builds on10
- TEACh: Task-Driven Embodied Agents That ChatAishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange et al.AAAI 2022 · 251 citations
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 228 citations
- FILM: Following Instructions in Language with Modular MethodsSo Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk et al.ICLR 2022 · 189 citations
- Mapping Natural Language Instructions to Mobile UI Action SequencesYang Li, Jiacong He, Xin Zhou, Yuan Zhang et al.ACL 2020 · 75 citations
- Factorizing Perception and Policy for Interactive Instruction FollowingKunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi et al.ICCV 2021 · 39 citations
Related papers
- Learning Compositional Tasks from Language InstructionsLajanugen Logeswaran, Wilka Carvalho, Honglak LeeAAAI 2023 · 4 citations
- Task Planning for Object Rearrangement in Multi-Room EnvironmentsKaran Mirakhor, Sourav Ghosh, Dipanjan Das, Brojeshwar BhowmickAAAI 2024 · 2 citations
- CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement LearningShun Qian, Bingquan Liu, Chengjie Sun, Peijin Xie et al.AAAI 2026
- Think before Go: Hierarchical Reasoning for Image-goal NavigationPengna Li, Kangyi Wu, Shaoqing Xu, Fang Li et al.ACL 2026 · 2 citations
- Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement LearningValerie Chen, Abhinav Gupta, Kenneth MarinoICLR 2021 · 6 citations
