Episodic Memory Question Answering
Samyak Datta, Sameer Dharur, Vincent Cartillier, Ruta Desai, Mukul Khanna, Dhruv Batra, Devi Parikh
摘要
Egocentric augmented reality devices such as wearable glasses passively capture visual data as a human wearer tours a home environment. We envision a scenario wherein the human communicates with an AI agent powering such a device by asking questions (e.g., “where did you last see my keys?”). In order to succeed at this task, the egocentric AI assistant must (1) construct semantically rich and efficient scene memories that encode spatio-temporal infor-mation about objects seen during the tour and (2) possess the ability to understand the question and ground its answer into the semantic memory representation. Towards that end, we introduce (1) a new task - Episodic Memory Question Answering (EMQA) wherein an egocentric AI assistant is provided with a video sequence (the tour) and a question as an input and is asked to localize its answer to the question within the tour, (2) a dataset of grounded questions designed to probe the agent's spatio-temporal understanding of the tour, and (3) a model for the task that encodes the scene as an allocentric, top-down semantic feature map and grounds the question into the map to localize the answer. We show that our choice of episodic scene memory outperforms naive, off-the-shelf solutions for the task as well as a host of very competitive baselines and is robust to noise in depth, pose as well as camera jitter.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu 等ICCV 2023 · 被引用 136 次
- VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of ThoughtGabriel Sarch, Lawrence Jang, Michael J. Tarr, William W. Cohen 等NeurIPS 2024 · 被引用 64 次
- EgoDistill: Egocentric Head Motion Distillation for Efficient Video UnderstandingShuhan Tan, Tushar Nagarajan, Kristen GraumanNeurIPS 2023 · 被引用 44 次
- EgoEnv: Human-centric environment representations from egocentric videoTushar Nagarajan, Santhosh Kumar Ramakrishnan, Ruta Desai, James Hillis 等NeurIPS 2023 · 被引用 28 次
- RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4DShuhei Kurita, Naoki Katsura, Eri OnamiICCV 2023 · 被引用 26 次
它引用的顶会 Paper8
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Bayesian Relational Memory for Semantic Visual NavigationYi Wu, Yuxin Wu, Aviv Tamar, Stuart Russell 等ICCV 2019 · 被引用 114 次
相关 Paper
- Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric ViewsVincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee 等AAAI 2021 · 被引用 89 次
- Agentic Very Long Video UnderstandingAniket Rege, Arka Sadhu, Yuliang Li, Kejie Li 等ACL 2026 · 被引用 8 次
- SpotEM: Efficient Video Search for Episodic MemorySanthosh Kumar Ramakrishnan, Ziad Al-Halah, Kristen GraumanICML 2023 · 被引用 15 次
- OpenEQA: Embodied Question Answering in the Era of Foundation ModelsArjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta 等CVPR 2024 · 被引用 44 次
- Ego-Grounding for Personalized Question-Answering in Egocentric VideosJunbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela YaoCVPR 2026 · 被引用 7 次
