Helping Hands: An Object-Aware Ego-Centric Video Recognition Model
Chuhan Zhang, Ankush Gupta, Andrew Zisserman
摘要
We introduce an object-aware decoder for improving the performance of spatio-temporal representations on egocentric videos. The key idea is to enhance object-awareness during training by tasking the model to predict hand positions, object positions, and the semantic label of the objects using paired captions when available. At inference time the model only requires RGB frames as inputs, and is able to track and ground objects (although it has not been trained explicitly for this). We demonstrate the performance of the object-aware representations learnt by our model, by: (i) evaluating it for strong transfer, i.e. through zero-shot testing, on a number of downstream video-text retrieval and classification benchmarks; and (ii) by using the representations learned as input for long-term video understanding tasks (e.g. Episodic Memory in Ego4D). In all cases the performance improves over the state of the art-even compared to networks trained with far larger batch sizes. We also show that by using noisy image-level detection as pseudo-labels in training, the model learns to provide better bounding boxes using video consistency, as well as grounding the words in the associated text descriptions. Overall, we show that the model can act as a drop-in replacement for an ego-centric video model to improve performance through visual-text grounding 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Multi-Factor Adaptive Vision Selection for Egocentric Video Question AnsweringHaoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song 等ICML 2024 · 被引用 23 次
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoTBaoqi Pei, Yifei Huang, Jilan Xu, Yuping He 等NeurIPS 2025 · 被引用 21 次
- Retrieval-Augmented Egocentric Video CaptioningJilan Xu, Yifei Huang, Junlin Hou, Guo Chen 等CVPR 2024 · 被引用 16 次
- OpenMMEgo: Enhancing Egocentric Understanding for LMMs with Open Weights and DataHao Luo, Zihao Yue, Wanpeng Zhang, Yicheng Feng 等NeurIPS 2025 · 被引用 10 次
- HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language ModelKhoa Vo, Thinh Phan, Kashu Yamazaki, Minh Tran 等NeurIPS 2024 · 被引用 8 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
相关 Paper
- Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation LearningBaoqi Pei, Yifei Huang, Jilan Xu, Guo Chen 等ICLR 2025
- EgoDTM: Towards 3D-Aware Egocentric Video-Language PretrainingBoshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng 等NeurIPS 2025 · 被引用 6 次
- Ego-Only: Egocentric Action Detection without Exocentric TransferringHuiyu Wang, Mitesh Kumar Singh, Lorenzo TorresaniICCV 2023 · 被引用 41 次
- HierVL: Learning Hierarchical Video-Language EmbeddingsKumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen GraumanCVPR 2023
- Single-Stage Visual Query Localization in Egocentric VideosHanwen Jiang, Santhosh Kumar Ramakrishnan, Kristen GraumanNeurIPS 2023 · 被引用 27 次
