Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents
Byeonghwi Kim, Jinyeon Kim, Yuyeong Kim, Cheolhong Min, Jonghyun Choi
Abstract
Accomplishing household tasks requires to plan step-by-step actions considering the consequences of previous actions. However, the state-of-the-art embodied agents often make mistakes in navigating the environment and interacting with proper objects due to imperfect learning by imitating experts or algorithmic planners without such knowledge. To improve both visual navigation and object interaction, we propose to consider the consequence of taken actions by CAPEAM (Context-Aware Planning and Environment-Aware Memory) that incorporates semantic context (e.g., appropriate objects to interact with) in a sequence of actions, and the changed spatial arrangement and states of interacted objects (e.g., location that the object has been moved to) in inferring the subsequent actions. We empirically show that the agent with the proposed CAPEAM achieves state-of-the-art performance in various metrics using a challenging interactive instruction following benchmark in both seen and unseen environments by large margins (up to +10.70% in unseen env.).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5df7d8a0-f82d-4713-8ec5-d9129decd138Cited by top-tier papers15
- Online Continual Learning for Interactive Instruction Following AgentsByeonghwi Kim, Minhyuk Seo, Jonghyun ChoiICLR 2024 · 22 citations
- Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory UtilizationTaeyoon Kwon, Dongwook Choi, Hyojun Kim, Sunghwan Kim et al.ICLR 2026 · 14 citations
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual ReasoningKailing Li, Qi'ao Xu, Tianwen Qian, Yuqian Fu et al.CVPR 2026 · 12 citations
- Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few ExamplesTaewoong Kim, Byeonghwi Kim, Jonghyun ChoiAAAI 2025 · 8 citations
- RoboAgent: Chaining Basic Capabilities for Embodied Task PlanningPeiran Xu, Jiaqi Zheng, Yadong MuCVPR 2026 · 6 citations
Builds on12
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- TEACh: Task-Driven Embodied Agents That ChatAishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange et al.AAAI 2022 · 251 citations
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 228 citations
- FILM: Following Instructions in Language with Modular MethodsSo Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk et al.ICLR 2022 · 189 citations
- Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language NavigationPeihao Chen, Dongyu Ji, Kunyang Lin, Runhao Zeng et al.NeurIPS 2022 · 143 citations
Related papers
- ACKnowledge: A Computational Framework for Human Compatible Affordance-based Interaction Planning in Real-world ContextsZiqi Pan, Xiucheng Zhang, Zisu Li, Zhenhui Peng et al.CHI 2025 · 1 citation
- Task Planning for Object Rearrangement in Multi-Room EnvironmentsKaran Mirakhor, Sourav Ghosh, Dipanjan Das, Brojeshwar BhowmickAAAI 2024 · 2 citations
- Bayesian Relational Memory for Semantic Visual NavigationYi Wu, Yuxin Wu, Aviv Tamar, Stuart Russell et al.ICCV 2019 · 114 citations
- Procedural Mistake Detection via Action Effect ModelingWenliang Guo, Yujiang Pu, Yu KongICLR 2026 · 6 citations
- Embodied Amodal Recognition: Learning to Move to Perceive ObjectsJianwei Yang, Zhile Ren, Mingze Xu, Xinlei Chen et al.ICCV 2019 · 70 citations
