Decoupling Dense Video Captioning via Task-specific Prompts
Wei Chen, Jianwei Niu, Xuefeng Liu, Xinghao Wu
摘要
Dense video captioning aims to generate descriptive sentences for each temporally localized event in a video. This task comprises two subtasks: event detection and event captioning. Existing methods commonly adopt a DETR-like (Detection Transformer) architecture to perform both subtasks in parallel. These methods assume that both subtasks require the same visual information and thus extract a single event representation for each event using a shared query. We observe that event detection and event captioning emphasize different regions of a video. In particular, compared to event captioning, event detection tends to focus more on the boundary regions of event proposals. Therefore, relying on shared queries may hinder the ability of the model to meet the specific needs of each subtask, leading to suboptimal performance. In this paper, we propose decoupling the two subtasks by assigning distinct queries to each, enabling more accurate capture of task-specific features. Specifically, we introduce a task-specific query transformation module. This module utilizes two sets of task-specific prompts to transform shared queries into queries tailored for each subtask. These task-specific queries enable each subtask to attend to the video regions that are most beneficial to its respective objectives. By integrating our method into several state-of-the-art frameworks, we achieve superior performance on both event detection and event captioning.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Task-Specific Information Decomposition for End-to-End Dense Video CaptioningZhiyue Liu, Xinru Zhang, Jinyuan LiuACL 2025
- End-to-End 3D Dense Captioning with Vote2Cap-DETRSijin Chen, Hongyuan Zhu, Xin Chen, Yinjie Lei 等CVPR 2023
- Sketch, Ground, and Refine: Top-Down Dense Video CaptioningChaorui Deng, Shizhe Chen, Da Chen, Yuan He 等CVPR 2021
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng 等ICCV 2021 · 被引用 238 次
- Towards Bridging Event Captioner and Sentence Localizer for Weakly Supervised Dense Event CaptioningShaoxiang Chen, Yu-Gang JiangCVPR 2021
