Cross-modal Fusion Transformer for Integrating Retrieved Knowledge into Video Caption Generation
Karina Abubakirova, Waseem Ullah, Latif U. Khan, Mohsen Guizani
摘要
In recent years, long-form video captioning has been an important task for indexing, accessibility, and downstream video analytics. However, it becomes challenging when videos are untrimmed and extend over minutes or hours. In these settings, visual evidence is often incomplete: objects are occluded, viewpoints shift, and crucial semantics may be implied rather than directly observable. Although recent hierarchical captioning models improve the temporal coverage, they still rely primarily on the available visual stream and can produce captions that omit key entities or relations. External knowledge retrieved from training data or knowledge bases can improve the visual signal, but existing integrations are often shallow and are either added late at the decoder input or applied without accounting for retrieval reliability, making them sensitive to noisy evidence and limiting cross-modal interaction. To solve this issue, we propose a video-aware knowledge fusion approach that integrates retrieved knowledge at the feature level prior to generation. The method combines (i) a contextual gating mechanism that modulates knowledge strength using retrieval confidence signals together with pooled visual context, and (ii) a cross-modal fusion transformer that refines the joint sequence of video and gated-knowledge query tokens through self-attention before decoding. We evaluate on challenging datasets such as YouCook2 and Ego4D-HCap, where our full model improves CIDEr by up to +11.1% on YouCook2 and by +4.2% / +5.8% on Ego4D-HCap segment descriptions and video summaries, compared with a strong baseline. This indicates that feature-level knowledge fusion can enhance semantic coverage in long-form caption generation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Text with Knowledge Graph Augmented Transformer for Video CaptioningXin Gu, Guang Chen, Yufei Wang, Libo Zhang 等CVPR 2023
- Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionMingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li 等AAAI 2026 · 被引用 1 次
- Hierarchical Context-aware Network for Dense Video Event CaptioningLei Ji, Xianglin Guo, Haoyang Huang, Xilin ChenACL 2021
- Do You Remember? Dense Video Captioning with Cross-Modal Memory RetrievalMinkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 等CVPR 2024
- Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image CaptioningXiaoxing You, Qiang Huang, Lingyu Li, Chi Zhang 等AAAI 2026 · 被引用 1 次
