Factorizing Perception and Policy for Interactive Instruction Following
Kunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi, Jonghyun Choi
Abstract
Performing simple household tasks based on language directives is very natural to humans, yet it remains an open challenge for AI agents. The ‘interactive instruction following’ task attempts to make progress towards building agents that jointly navigate, interact, and reason in the environment at every step. To address the multifaceted problem, we propose a model that factorizes the task into interactive perception and action policy streams with enhanced components and name it as MOCA, a Modular Object-Centric Approach. We empirically validate that MOCA outperforms prior arts by significant margins on the ALFRED benchmark with improved generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30df37b5-2082-4018-b935-3d9770c21cb6Cited by top-tier papers12
- Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied AgentsByeonghwi Kim, Jinyeon Kim, Yuyeong Kim, Cheolhong Min et al.ICCV 2023 · 46 citations
- Online Continual Learning for Interactive Instruction Following AgentsByeonghwi Kim, Minhyuk Seo, Jonghyun ChoiICLR 2024 · 22 citations
- One Step at a Time: Long-Horizon Vision-and-Language Navigation with MilestonesChan Hee Song, Jihyung Kil, Tai-Yu Pan, Brian M. Sadler et al.CVPR 2022 · 21 citations
- Multi-Level Compositional Reasoning for Interactive Instruction FollowingSuvaansh Bhambri, Byeonghwi Kim, Jonghyun ChoiAAAI 2023 · 14 citations
- Guardian: A Runtime Framework for LLM-Based UI ExplorationDezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu et al.ISSTA 2024 · 13 citations
Builds on4
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- GridToPix: Training Embodied Agents with Minimal SupervisionUnnat Jain, Iou-Jen Liu, Svetlana Lazebnik, Aniruddha Kembhavi et al.ICCV 2021 · 25 citations
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk et al.CVPR 2020
Related papers
- Egocentric Planning for Scalable Embodied Task AchievementXiaotian Liu, Héctor Palacios, Christian MuiseNeurIPS 2023 · 9 citations
- CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-AffordanceJinming Li, Yichen Zhu, Zhibin Tang, Junjie Wen et al.ICCV 2025 · 7 citations
- Human-Object Interaction from Human-level InstructionsZhen Wu, Jiaman Li, Pei Xu, C. Karen LiuICCV 2025 · 3 citations
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu et al.ICCV 2025 · 2 citations
- Ins-DetCLIP: Aligning Detection Model to Follow Human-Language InstructionRenjie Pi, Lewei Yao, Jianhua Han, Xiaodan Liang et al.ICLR 2024 · 5 citations
