Grounded Question-Answering in Long Egocentric Videos
Shangzhe Di, Weidi Xie
Abstract
Existing approaches to video understanding, mainly designed for short videos from a third-person perspective, are limited in their applicability in certain fields, such as robotics. In this paper, we delve into open-ended questionanswering (QA) in long, egocentric videos, which allows individuals or robots to inquire about their own past visual experiences. This task presents unique challenges, including the complexity of temporally grounding queries within extensive video content, the high resource demands for precise data annotation, and the inherent difficulty of evaluating open-ended answers due to their ambiguous nature. Our proposed approach tackles these challenges by (i) integrating query grounding and answering within a unified model to reduce error propagation; (ii) employing large language models for efficient and scalable data synthesis; and (iii) introducing a close-ended QA task for evaluation, to manage answer ambiguity. Extensive experiments demonstrate the effectiveness of our method, which also achieves stateof-the-art performance on the QAEgo4D and Ego4D-NLQ benchmarks. Code, data, and models are open-sourced 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2128c3ed-e5c3-43af-914a-6ca0506d8681Cited by top-tier papers31
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming AssistantHaibo Wang, Bo Feng, Zhengfeng Lai, Mingze Xu et al.NeurIPS 2025 · 63 citations
- ViSpeak: Visual Instruction Feedback in Streaming VideosShenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng et al.ICCV 2025 · 42 citations
- VideoITG: Multimodal Video Understanding with Instructed Temporal GroundingShihao Wang, Guo Chen, De-An Huang, Zhiqi Li et al.CVPR 2026 · 35 citations
- Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosQirui Chen, Shangzhe Di, Weidi XieAAAI 2025 · 35 citations
- Streaming Video Instruction TuningJiaer Xia, Peixian Chen, Mengdan Zhang, Xing Sun et al.CVPR 2026 · 28 citations
Builds on18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li et al.AAAI 2020 · 4,823 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- Ego-Grounding for Personalized Question-Answering in Egocentric VideosJunbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela YaoCVPR 2026 · 7 citations
- ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition BenchmarkRonghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin et al.CVPR 2025
- MMEgo: Towards Building Egocentric Multimodal LLMs for Video QAHanrong Ye, Haotian Zhang, Erik A. Daxberger, Lin Chen et al.ICLR 2025
- EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language ModelsSijie Cheng, Zhicheng Guo, Jingwen Wu, Kechen Fang et al.CVPR 2024
- NaQ: Leveraging Narrations as Queries to Supervise Episodic MemorySanthosh Kumar Ramakrishnan, Ziad Al-Halah, Kristen GraumanCVPR 2023
