Lune

KDD2026顶会

Cross-modal Fusion Transformer for Integrating Retrieved Knowledge into Video Caption Generation

Karina Abubakirova, Waseem Ullah, Latif U. Khan, Mohsen Guizani

2026年份

摘要

In recent years, long-form video captioning has been an important task for indexing, accessibility, and downstream video analytics. However, it becomes challenging when videos are untrimmed and extend over minutes or hours. In these settings, visual evidence is often incomplete: objects are occluded, viewpoints shift, and crucial semantics may be implied rather than directly observable. Although recent hierarchical captioning models improve the temporal coverage, they still rely primarily on the available visual stream and can produce captions that omit key entities or relations. External knowledge retrieved from training data or knowledge bases can improve the visual signal, but existing integrations are often shallow and are either added late at the decoder input or applied without accounting for retrieval reliability, making them sensitive to noisy evidence and limiting cross-modal interaction. To solve this issue, we propose a video-aware knowledge fusion approach that integrates retrieved knowledge at the feature level prior to generation. The method combines (i) a contextual gating mechanism that modulates knowledge strength using retrieval confidence signals together with pooled visual context, and (ii) a cross-modal fusion transformer that refines the joint sequence of video and gated-knowledge query tokens through self-attention before decoding. We evaluate on challenging datasets such as YouCook2 and Ego4D-HCap, where our full model improves CIDEr by up to +11.1% on YouCook2 and by +4.2% / +5.8% on Ego4D-HCap segment descriptions and video summaries, compared with a strong baseline. This indicates that feature-level knowledge fusion can enhance semantic coverage in long-form caption generation.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 65b7cf8a-ac0e-40c6-baac-6e4e0bf704cc

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖