An Efficient Framework for Dense Video Captioning
Maitreya Suin, A. N. Rajagopalan
摘要
Dense video captioning is an extremely challenging task since an accurate and faithful description of events in a video requires a holistic knowledge of the video contents as well as contextual reasoning of individual events. Most existing approaches handle this problem by first proposing event boundaries from a video and then captioning on a subset of the proposals. Generation of dense temporal annotations and corresponding captions from long videos can be dramatically source consuming. In this paper, we focus on the task of generating a dense description of temporally untrimmed videos and aim to significantly reduce the computational cost by processing fewer frames while maintaining accuracy. Existing video captioning methods sample frames with a predefined frequency over the entire video or use all the frames. Instead, we propose a deep reinforcement-based approach which enables an agent to describe multiple events in a video by watching a portion of the frames. The agent needs to watch more frames when it is processing an informative part of the video, and skip frames when there is redundancy. The agent is trained using actor-critic algorithm, where the actor determines the frames to be watched from a video and the critic assesses the optimality of the decisions taken by the actor. Such an efficient frame selection simplifies the event proposal task considerably. This has the added effect of reducing the occurrence of unwanted proposals. The encoded state representation of the frame selection agent is further utilized for guiding event proposal and caption generation tasks. We also leverage the idea of knowledge distillation to improve the accuracy. We conduct extensive evaluations on ActivityNet captions dataset to validate our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng 等ICCV 2021 · 被引用 238 次
- Distillation-guided Image InpaintingMaitreya Suin, Kuldeep Purohit, A. N. RajagopalanICCV 2021 · 被引用 42 次
- Exploiting Auxiliary Caption for Video GroundingHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li 等AAAI 2024 · 被引用 16 次
- SpotEM: Efficient Video Search for Episodic MemorySanthosh Kumar Ramakrishnan, Ziad Al-Halah, Kristen GraumanICML 2023 · 被引用 15 次
- Token Mixing: Parameter-Efficient Transfer Learning from Image-Language to Video-LanguageYuqi Liu, Luhui Xu, Pengfei Xiong, Qin JinAAAI 2023 · 被引用 10 次
相关 Paper
- Towards Bridging Event Captioner and Sentence Localizer for Weakly Supervised Dense Event CaptioningShaoxiang Chen, Yu-Gang JiangCVPR 2021
- Do You Remember? Dense Video Captioning with Cross-Modal Memory RetrievalMinkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 等CVPR 2024
- Sketch, Ground, and Refine: Top-Down Dense Video CaptioningChaorui Deng, Shizhe Chen, Da Chen, Yuan He 等CVPR 2021
- Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionMingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li 等AAAI 2026 · 被引用 1 次
- Event-Equalized Dense Video CaptioningKangyi Wu, Pengna Li, Jingwen Fu, Yizhe Li 等CVPR 2025
