Factorizing Perception and Policy for Interactive Instruction Following
Kunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi, Jonghyun Choi
摘要
Performing simple household tasks based on language directives is very natural to humans, yet it remains an open challenge for AI agents. The ‘interactive instruction following’ task attempts to make progress towards building agents that jointly navigate, interact, and reason in the environment at every step. To address the multifaceted problem, we propose a model that factorizes the task into interactive perception and action policy streams with enhanced components and name it as MOCA, a Modular Object-Centric Approach. We empirically validate that MOCA outperforms prior arts by significant margins on the ALFRED benchmark with improved generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied AgentsByeonghwi Kim, Jinyeon Kim, Yuyeong Kim, Cheolhong Min 等ICCV 2023 · 被引用 46 次
- Online Continual Learning for Interactive Instruction Following AgentsByeonghwi Kim, Minhyuk Seo, Jonghyun ChoiICLR 2024 · 被引用 22 次
- One Step at a Time: Long-Horizon Vision-and-Language Navigation with MilestonesChan Hee Song, Jihyung Kil, Tai-Yu Pan, Brian M. Sadler 等CVPR 2022 · 被引用 21 次
- Multi-Level Compositional Reasoning for Interactive Instruction FollowingSuvaansh Bhambri, Byeonghwi Kim, Jonghyun ChoiAAAI 2023 · 被引用 14 次
- Guardian: A Runtime Framework for LLM-Based UI ExplorationDezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu 等ISSTA 2024 · 被引用 13 次
它引用的顶会 Paper4
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- GridToPix: Training Embodied Agents with Minimal SupervisionUnnat Jain, Iou-Jen Liu, Svetlana Lazebnik, Aniruddha Kembhavi 等ICCV 2021 · 被引用 25 次
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk 等CVPR 2020
相关 Paper
- Egocentric Planning for Scalable Embodied Task AchievementXiaotian Liu, Héctor Palacios, Christian MuiseNeurIPS 2023 · 被引用 9 次
- CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-AffordanceJinming Li, Yichen Zhu, Zhibin Tang, Junjie Wen 等ICCV 2025 · 被引用 7 次
- Human-Object Interaction from Human-level InstructionsZhen Wu, Jiaman Li, Pei Xu, C. Karen LiuICCV 2025 · 被引用 3 次
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu 等ICCV 2025 · 被引用 2 次
- Ins-DetCLIP: Aligning Detection Model to Follow Human-Language InstructionRenjie Pi, Lewei Yao, Jianhua Han, Xiaodan Liang 等ICLR 2024 · 被引用 5 次
