Query-Focused Multimodal Summarization with Gate-Guided Mixture-of-Experts
Jiajun Han, Xuran Yang, Hui Zhang
摘要
The goal of generic multimodal summarization is to extract the most important information from different modalities to form summaries. Yet the importance of scenes and text in a video is often subjective, and users should have the option of customizing the summary by using natural language to specify what is important to them. However, existing methods for fully automatic multimodal summarization have not exploited available language models, which can serve as an effective prior for saliency. To address this issue, we introduce Query-Focused Multimodal Summ arization(QFSumm), a single framework for addressing both generic and query-focused multimodal summarization, typically approached separately in the literature. In addition, we propose a novel gate-guided mixture-of-experts that uses expert gate module to organize three experts (video expert, text expert and shared expert) to model the correlations between multimodal information. In addition, we propose two novel contrastive losses to represent consistency and diversity. Extensive experiments on a query-focused video summarization dataset (QFVS), two standard video summarization datasets (TVSum and SumMe) and three multimodal summarization datasets (CNN, Daily Mail and BLiSS) demonstrate the superiority of QFSumm, achieving state-of-the-art performances on all datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 被引用 196 次
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui 等CVPR 2023
- IntentVizor: Towards Generic Query Guided Interactive Video SummarizationGuande Wu, Jianzhe Lin, Cláudio T. SilvaCVPR 2022 · 被引用 36 次
- SD-VSum: A Method and Dataset for Script-Driven Video SummarizationManolis Mylonas, Evlampios Apostolidis, Vasileios MezarisACM MM 2025 · 被引用 2 次
- TripleSumm: Adaptive Triple-Modality Fusion for Video SummarizationSumin Kim, Hyemin Jeong, Mingu Kang, Yejin Kim 等ICLR 2026 · 被引用 2 次
