Retrieval-Augmented Egocentric Video Captioning
Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, Weidi Xie
Abstract
Understanding human actions from videos offirst-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper, (1) we develop EgoInstructor, a retrieval-augmented multimodal captioning model that automatically retrieves semantically relevant third-person instructional videos to enhance the video captioning of egocentric videos, (2) for training the cross-view retrieval module, we devise an au-tomatic pipeline to discover ego-exo video pairs from distinct large-scale egocentric and exocentric datasets, (3) we train the cross-view retrieval module with a novel EgoEx-oNCE loss that pulls egocentric and exocentric video features closer, by aligning them to shared text features that describe similar actions, (4) through extensive experiments, our cross-view retrieval module demonstrates superior performance across seven benchmarks. Regarding egocen-tric video captioning, EgoInstructor exhibits significant improvements by leveraging third-person videos as references.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7fb2b70-c441-44c9-95f1-539124a07f7eCited by top-tier papers28
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoTBaoqi Pei, Yifei Huang, Jilan Xu, Yuping He et al.NeurIPS 2025 · 21 citations
- Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language GeneratorRonglai Zuo, Rolandos Alexandros Potamias, Evangelos Ververas, Jiankang Deng et al.ICCV 2025 · 9 citations
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMsInsu Lee, Wooje Park, Jaeyun Jang, Minyoung Noh et al.NeurIPS 2025 · 8 citations
- Tell, Don't Show: Language Guidance Eases Transfer Across Domains in Images and VideosTarun Kalluri, Bodhisattwa Prasad Majumder, Manmohan ChandrakerICML 2024 · 7 citations
- Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue ConsistencyZhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu et al.CVPR 2026 · 3 citations
Builds on35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
Related papers
- Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video RepresentationsJungin Park, Jiyoung Lee, Kwanghoon SohnCVPR 2025
- Ego-Exo: Transferring Visual Representations From Third-Person to First-Person VideosYanghao Li, Tushar Nagarajan, Bo Xiong, Kristen GraumanCVPR 2021
- Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video UnderstandingHaoyu Zhang, Qiaohui Chu, Meng Liu, Haoxiang Shi et al.AAAI 2026 · 17 citations
- Ego-Only: Egocentric Action Detection without Exocentric TransferringHuiyu Wang, Mitesh Kumar Singh, Lorenzo TorresaniICCV 2023 · 41 citations
- Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation LearningMi Luo, Zihui Xue, Alex Dimakis, Kristen GraumanCVPR 2025
