Egocentric Planning for Scalable Embodied Task Achievement
Xiaotian Liu, Héctor Palacios, Christian Muise
摘要
Embodied agents face significant challenges when tasked with performing actions in diverse environments, particularly in generalizing across object types and executing suitable actions to accomplish tasks. Furthermore, agents should exhibit robustness, minimizing the execution of illegal actions. In this work, we present Egocentric Planning, an innovative approach that combines symbolic planning and Object-oriented POMDPs to solve tasks in complex environments, harnessing existing models for visual perception and natural language processing. We evaluated our approach in ALFRED, a simulated environment designed for domestic tasks, and demonstrated its high scalability, achieving an impressive 36.07% unseen success rate in the ALFRED benchmark and winning the ALFRED challenge at CVPR Embodied AI workshop. Our method requires reliable perception and the specification or learning of a symbolic description of the preconditions and effects of the agent's actions, as well as what object types reveal information about others. It is capable of naturally scaling to solve new tasks beyond ALFRED, as long as they can be solved using the available skills. This work offers a solid baseline for studying end-to-end and hybrid methods that aim to generalize to new tasks, including recent approaches relying on LLMs, but often struggle to scale to long sequences of actions or produce robust plans for novel tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MO-DDN: A Coarse-to-Fine Attribute-based Exploration Agent for Multi-Object Demand-driven NavigationHongcheng Wang, Peiqi Liu, Wenzhe Cai, Mingdong Wu 等NeurIPS 2024 · 被引用 12 次
- Graph Learning for Numeric PlanningDillon Z. Chen, Sylvie ThiébauxNeurIPS 2024 · 被引用 8 次
- EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real WorldYifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang 等CVPR 2024
- Minerva-Ego: Spatiotemporal Hints for Egocentric Video UnderstandingArsha Nagrani, Jasper Uijlings, Shyamal Buch, Tobias Weyand 等CVPR 2026
- Graph-Theoretic Intrinsic Reward: Guiding RL with Effective ResistanceJatin Chauhan, Shivam Bhardwaj, Aditya Saibewar, Aditya Ramesh 等ICLR 2026
它引用的顶会 Paper7
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- FILM: Following Instructions in Language with Modular MethodsSo Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk 等ICLR 2022 · 被引用 189 次
- Continual Learning via Local Module CompositionOleksiy Ostapenko, Pau Rodríguez, Massimo Caccia, Laurent CharlinNeurIPS 2021 · 被引用 98 次
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 被引用 51 次
- Visual Room RearrangementLuca Weihs, Matt Deitke, Aniruddha Kembhavi, Roozbeh MottaghiCVPR 2021
相关 Paper
- Factorizing Perception and Policy for Interactive Instruction FollowingKunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi 等ICCV 2021 · 被引用 39 次
- Natural Language PDDL (NL-PDDL) for Open-world Goal-oriented Commonsense Regression Planning in Embodied AIXiaotian Liu, Armin Toroghi, Jiazhou Liang, David Courtis 等ICLR 2026
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk 等CVPR 2020
- R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action PlannerZiyi Bai, Hanxuan Li, Bin Fu, Chuyan Xiong 等CVPR 2025
- EPO: Hierarchical LLM Agents with Environment Preference OptimizationQi Zhao, Haotian Fu, Chen Sun, George KonidarisEMNLP 2024 · 被引用 3 次
