SpotEM: Efficient Video Search for Episodic Memory
Santhosh Kumar Ramakrishnan, Ziad Al-Halah, Kristen Grauman
摘要
The goal in episodic memory (EM) is to search a long egocentric video to answer a natural language query (e.g.,"where did I leave my purse?"). Existing EM methods exhaustively extract expensive fixed-length clip features to look everywhere in the video for the answer, which is infeasible for long wearable-camera videos that span hours or even days. We propose SpotEM, an approach to achieve efficiency for a given EM method while maintaining good accuracy. SpotEM consists of three key ideas: 1) a novel clip selector that learns to identify promising video regions to search conditioned on the language query; 2) a set of low-cost semantic indexing features that capture the context of rooms, objects, and interactions that suggest where to look; and 3) distillation losses that address the optimization issues arising from end-to-end joint training of the clip selector and EM model. Our experiments on 200+ hours of video from the Ego4D EM Natural Language Queries benchmark and three different EM models demonstrate the effectiveness of our approach: computing only 10% - 25% of the clip features, we preserve 84% - 97% of the original EM model's accuracy. Project page: https://vision.cs.utexas.edu/projects/spotem
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou 等AAAI 2024 · 被引用 30 次
- Multi-Factor Adaptive Vision Selection for Egocentric Video Question AnsweringHaoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song 等ICML 2024 · 被引用 23 次
- HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba PoolingJoungbin An, Kristen GraumanCVPR 2026 · 被引用 2 次
- A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task PerspectivesSimone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, Giuseppe AvertaCVPR 2024 · 被引用 1 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 被引用 1,226 次
相关 Paper
- Episodic Memory Question AnsweringSamyak Datta, Sameer Dharur, Vincent Cartillier, Ruta Desai 等CVPR 2022 · 被引用 23 次
- Interactive Episodic Memory with User FeedbackNikesh Subedi, Loris Bazzani, Ziad Al-HalahCVPR 2026
- Single-Stage Visual Query Localization in Egocentric VideosHanwen Jiang, Santhosh Kumar Ramakrishnan, Kristen GraumanNeurIPS 2023 · 被引用 27 次
- ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental LearningJongseo Lee, Kyungho Bae, Kyle Min, Gyeong-Moon Park 等ICCV 2025 · 被引用 2 次
- EgoDistill: Egocentric Head Motion Distillation for Efficient Video UnderstandingShuhan Tan, Tushar Nagarajan, Kristen GraumanNeurIPS 2023 · 被引用 44 次
