Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
Yue Fan, Xiaojian Ma, Rongpeng Su, Jun Guo, Rujie Wu, Xi Chen, Qing Li
Abstract
This paper investigates the problem of understanding dynamic 3D scenes from egocentric observations, a key challenge in robotics and embodied AI. Unlike prior studies that explored this as long-form video understanding and utilized egocentric video only, we instead propose an LLMbased agent, Embodied VideoAgent, which constructs scene memory from both egocentric video and embodied sensory inputs (e.g. depth and pose sensing). We further introduce a VLM-based approach to automatically update the memory when actions or activities over objects are perceived. Embodied VideoAgent attains significant advantages over counterparts in challenging reasoning and planning tasks in 3D scenes, achieving gains of 6.5% on Ego4D-VQ3D, 2.6% on OpenEQA, and 15.3% on EnvQA. We have also demonstrated its potential in various embodied AI tasks including generating embodied interactions and perception for robot manipulation. The code and demo will be made public.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term MemoryLin Long, Yichen He, Wentao Ye, Yiyuan Pan et al.ICLR 2026 · 90 citations
- Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference TuningPengxiang Li, Zhi Gao, Bofei Zhang, Yapeng Mi et al.NeurIPS 2025 · 21 citations
- Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied NavigationZiyu Zhu, Xilin Wang, Yixuan Li, Zhuofan Zhang et al.ICCV 2025 · 11 citations
- D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and NavigationZihan Wang, Seungjun Lee, Guangzhao Dai, Gim Hee LeeCVPR 2026 · 9 citations
- STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement LearningXiaowen Zhang, Zhi Gao, Licheng Jiao, Lingling Li et al.ICLR 2026 · 3 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint FramesSahithya Ravi, Gabriel Herbert Sarch, Vibhav Vineet, Andrew D. Wilson et al.EMNLP 2025 · 1 citation
- Understanding Dynamic Scenes in Ego Centric 4D Point CloudsJunsheng Huang, Shengyu Hao, Bocheng Hu, Hongwei Wang et al.AAAI 2026 · 4 citations
- 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language ModelWenbo Hu, Yining Hong, Yanjun Wang, Leison Gao et al.NeurIPS 2025 · 30 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- Extending Embodied Question Answering from Perception to DecisionXicheng Gong, Qiwei Li, Peiran Xu, Yadong MuCVPR 2026 · 1 citation
