Egocentric Vehicle Dense Video Captioning
Feiyu Chen, Cong Xu, Qi Jia, Yihua Wang, Yuhan Liu, Haotian Zhang, Endong Wang
Abstract
Traditional dense video captioning predominantly focuses on edited exocentric footage. These videos are filmed from an external perspective and generally feature distinct transitions between different events, as exemplified in edited instructional videos. However, such videos do not genuinely reflect the way we perceive our real lives. Instead, we observe the world from an egocentric viewpoint and witness only continuous unedited footage. To facilitate further research, we introduce a new topic: Egocentric Vehicle Dense Video Captioning, in classic vehicle driving scenarios. This is a multi-modal, multi-task subject endeavor for a comprehensive understanding of untrimmed, egocentric driving videos. It consists of three sub-tasks that concentrate on event localization, captioning, and vehicle state estimation separately. To accomplish these tasks, it is necessary to deal with at least three challenges: extracting ego-motion relevant information, describing driving behavior and analyzing the underlying rationale, as well as resolving the boundary ambiguity problem. In response, we devise corresponding solutions, including a vehicle ego-motion learning strategy and a novel adjacent contrastive learning strategy, which effectively address the aforementioned issues. We validate our method by conducting extensive experiments on the BDD-X dataset, all of which show promising results and achieve new state-of-the-art performance on most metrics, which proves the effect of our approach.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers6
- Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus AdaptationZhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu et al.ACM MM 2025
- Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction TuningJi Hyeok Jung, Eun Tae Kim, Seo Yeon Kim, Joo Ho Lee et al.CVPR 2025
- MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video UnderstandingTongtong Cheng, Rongzhen Li, Yixin Xiong, Tao Zhang et al.ICCV 2025
- DMC3: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question AnsweringJiayi Zou, Chaofan Chen, Bing-Kun Bao, Changsheng XuACM MM 2025
- Self-Critical Distillation Network for Video-based Commonsense CaptioningMengqi Yuan, Gengyun Jia, Bing-Kun BaoCVPR 2026
Related papers
- Do You Remember? Dense Video Captioning with Cross-Modal Memory RetrievalMinkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi et al.CVPR 2024
- EgoDTM: Towards 3D-Aware Egocentric Video-Language PretrainingBoshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng et al.NeurIPS 2025 · 6 citations
- Task-Specific Information Decomposition for End-to-End Dense Video CaptioningZhiyue Liu, Xinru Zhang, Jinyuan LiuACL 2025
- Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-LearningZhuyang Xie, Yan Yang, Yankai Yu, Jie Wang et al.AAAI 2025 · 5 citations
- Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video CaptioningSeungHyup Baek, Jimin Lee, Hyeongkeun Lee, Jae Won ChoCVPR 2026 · 1 citation
