Decoupling Dense Video Captioning via Task-specific Prompts
Wei Chen, Jianwei Niu, Xuefeng Liu, Xinghao Wu
Abstract
Dense video captioning aims to generate descriptive sentences for each temporally localized event in a video. This task comprises two subtasks: event detection and event captioning. Existing methods commonly adopt a DETR-like (Detection Transformer) architecture to perform both subtasks in parallel. These methods assume that both subtasks require the same visual information and thus extract a single event representation for each event using a shared query. We observe that event detection and event captioning emphasize different regions of a video. In particular, compared to event captioning, event detection tends to focus more on the boundary regions of event proposals. Therefore, relying on shared queries may hinder the ability of the model to meet the specific needs of each subtask, leading to suboptimal performance. In this paper, we propose decoupling the two subtasks by assigning distinct queries to each, enabling more accurate capture of task-specific features. Specifically, we introduce a task-specific query transformation module. This module utilizes two sets of task-specific prompts to transform shared queries into queries tailored for each subtask. These task-specific queries enable each subtask to attend to the video regions that are most beneficial to its respective objectives. By integrating our method into several state-of-the-art frameworks, we achieve superior performance on both event detection and event captioning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d80c7bd5-818d-4d2e-af99-db85dd511b9bCited by top-tier papers1
Ask how each one uses itRelated papers
- Task-Specific Information Decomposition for End-to-End Dense Video CaptioningZhiyue Liu, Xinru Zhang, Jinyuan LiuACL 2025
- End-to-End 3D Dense Captioning with Vote2Cap-DETRSijin Chen, Hongyuan Zhu, Xin Chen, Yinjie Lei et al.CVPR 2023
- Sketch, Ground, and Refine: Top-Down Dense Video CaptioningChaorui Deng, Shizhe Chen, Da Chen, Yuan He et al.CVPR 2021
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng et al.ICCV 2021 · 238 citations
- Towards Bridging Event Captioner and Sentence Localizer for Weakly Supervised Dense Event CaptioningShaoxiang Chen, Yu-Gang JiangCVPR 2021
