Lune

ACM MM2025Top-tier venue

Query-Focused Multimodal Summarization with Gate-Guided Mixture-of-Experts

Jiajun Han, Xuran Yang, Hui Zhang

2025Year
1Citations

Abstract

The goal of generic multimodal summarization is to extract the most important information from different modalities to form summaries. Yet the importance of scenes and text in a video is often subjective, and users should have the option of customizing the summary by using natural language to specify what is important to them. However, existing methods for fully automatic multimodal summarization have not exploited available language models, which can serve as an effective prior for saliency. To address this issue, we introduce Query-Focused Multimodal Summ arization(QFSumm), a single framework for addressing both generic and query-focused multimodal summarization, typically approached separately in the literature. In addition, we propose a novel gate-guided mixture-of-experts that uses expert gate module to organize three experts (video expert, text expert and shared expert) to model the correlations between multimodal information. In addition, we propose two novel contrastive losses to represent consistency and diversity. Extensive experiments on a query-focused video summarization dataset (QFVS), two standard video summarization datasets (TVSum and SumMe) and three multimodal summarization datasets (CNN, Daily Mail and BLiSS) demonstrate the superiority of QFSumm, achieving state-of-the-art performances on all datasets.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 0078a110-0f03-48fc-9197-41b4bb9bed0a

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines