Egocentric Planning for Scalable Embodied Task Achievement
Xiaotian Liu, Héctor Palacios, Christian Muise
Abstract
Embodied agents face significant challenges when tasked with performing actions in diverse environments, particularly in generalizing across object types and executing suitable actions to accomplish tasks. Furthermore, agents should exhibit robustness, minimizing the execution of illegal actions. In this work, we present Egocentric Planning, an innovative approach that combines symbolic planning and Object-oriented POMDPs to solve tasks in complex environments, harnessing existing models for visual perception and natural language processing. We evaluated our approach in ALFRED, a simulated environment designed for domestic tasks, and demonstrated its high scalability, achieving an impressive 36.07% unseen success rate in the ALFRED benchmark and winning the ALFRED challenge at CVPR Embodied AI workshop. Our method requires reliable perception and the specification or learning of a symbolic description of the preconditions and effects of the agent's actions, as well as what object types reveal information about others. It is capable of naturally scaling to solve new tasks beyond ALFRED, as long as they can be solved using the available skills. This work offers a solid baseline for studying end-to-end and hybrid methods that aim to generalize to new tasks, including recent approaches relying on LLMs, but often struggle to scale to long sequences of actions or produce robust plans for novel tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 201525d9-53bd-4d29-a098-af50b7632015Cited by top-tier papers5
- MO-DDN: A Coarse-to-Fine Attribute-based Exploration Agent for Multi-Object Demand-driven NavigationHongcheng Wang, Peiqi Liu, Wenzhe Cai, Mingdong Wu et al.NeurIPS 2024 · 12 citations
- Graph Learning for Numeric PlanningDillon Z. Chen, Sylvie ThiébauxNeurIPS 2024 · 8 citations
- EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real WorldYifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang et al.CVPR 2024
- Minerva-Ego: Spatiotemporal Hints for Egocentric Video UnderstandingArsha Nagrani, Jasper Uijlings, Shyamal Buch, Tobias Weyand et al.CVPR 2026
- Graph-Theoretic Intrinsic Reward: Guiding RL with Effective ResistanceJatin Chauhan, Shivam Bhardwaj, Aditya Saibewar, Aditya Ramesh et al.ICLR 2026
Builds on7
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- FILM: Following Instructions in Language with Modular MethodsSo Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk et al.ICLR 2022 · 189 citations
- Continual Learning via Local Module CompositionOleksiy Ostapenko, Pau Rodríguez, Massimo Caccia, Laurent CharlinNeurIPS 2021 · 98 citations
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 51 citations
- Visual Room RearrangementLuca Weihs, Matt Deitke, Aniruddha Kembhavi, Roozbeh MottaghiCVPR 2021
Related papers
- Factorizing Perception and Policy for Interactive Instruction FollowingKunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi et al.ICCV 2021 · 39 citations
- Natural Language PDDL (NL-PDDL) for Open-world Goal-oriented Commonsense Regression Planning in Embodied AIXiaotian Liu, Armin Toroghi, Jiazhou Liang, David Courtis et al.ICLR 2026
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk et al.CVPR 2020
- R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action PlannerZiyi Bai, Hanxuan Li, Bin Fu, Chuyan Xiong et al.CVPR 2025
- EPO: Hierarchical LLM Agents with Environment Preference OptimizationQi Zhao, Haotian Fu, Chen Sun, George KonidarisEMNLP 2024 · 3 citations
