A Knowledge Augmented and Multimodal-Based Framework for Video Summarization
Jiehang Xie, Xuanbai Chen, Shao-Ping Lu, Yulu Yang
Abstract
Video summarization aims to generate a compact version of a lengthy video that retains its primary content. In general, humans are gifted with producing a high-quality video summary, because they acquire crucial content through multiple dimensional information and own abundant background knowledge about the original video. However, existing methods rarely consider multichannel information and ignore the impact of external knowledge, resulting in the limited quality of the generated summaries. This paper proposes a knowledge augmented and multimodal-based video summarization method, termed KAMV, to address the problem above. Specifically, we design a knowledge encoder with a hybrid method consisting of generation and retrieval, to capture descriptive content and latent connections between events and entities based on the external knowledge base, which can provide rich implicit knowledge for better comprehending the video viewed. Furthermore, for the sake of exploring the interactions among visual, audio, implicit knowledge and emphasizing the content that is most relevant to the desired summary, we present a fusion module under the supervision of these multimodal information. By conducting extensive experiments on four public datasets, the results demonstrate the superior performance yielded by the proposed KAMV compared to the state-of-the-art video summarization approaches.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 95d4c45b-dd83-4b51-b679-fe02dbfca620Cited by top-tier papers1
Ask how each one uses itRelated papers
- IntentVizor: Towards Generic Query Guided Interactive Video SummarizationGuande Wu, Jianzhe Lin, Cláudio T. SilvaCVPR 2022 · 36 citations
- Video Summarization with Large Language ModelsMin Jung Lee, Dayoung Gong, Minsu ChoCVPR 2025
- Query-Focused Multimodal Summarization with Gate-Guided Mixture-of-ExpertsJiajun Han, Xuran Yang, Hui ZhangACM MM 2025 · 1 citation
- Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 VideosNayu Liu, Kaiwen Wei, Xian Sun, Hongfeng Yu et al.EMNLP 2022 · 10 citations
- Cross-modal Fusion Transformer for Integrating Retrieved Knowledge into Video Caption GenerationKarina Abubakirova, Waseem Ullah, Latif U. Khan, Mohsen GuizaniKDD 2026
