Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain Videos
Nayu Liu, Xian Sun, Hongfeng Yu, Wenkai Zhang, Guangluan Xu
Abstract
Multimodal summarization for open-domain videos is an emerging task, aiming to generate a summary from multisource information (video, audio, transcript). Despite the success of recent multiencoder-decoder frameworks on this task, existing methods lack finegrained multimodality interactions of multisource inputs. Besides, unlike other multimodal tasks, this task has longer multimodal sequences with more redundancy and noise. To address these two issues, we propose a multistage fusion network with the fusion forget gate module, which builds upon this approach by modeling fine-grained interactions between the multisource modalities through a multistep fusion schema and controlling the flow of redundant information between multimodal long sequences via a forgetting module. Experimental results on the How2 dataset show that our proposed model achieves a new state-of-the-art performance. Comprehensive analysis empirically verifies the effectiveness of our fusion schema and forgetting module on multiple encoder-decoder architectures. Specially, when using high noise ASR transcripts (W ER>30%), our model still achieves performance close to the ground-truth transcript model, which reduces manual annotation cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ffeedc7-3ea6-4907-b3f6-5d9ba7e5c88bCited by top-tier papers12
- Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive SummarizationTiezheng Yu, Wenliang Dai, Zihan Liu, Pascale FungEMNLP 2021 · 64 citations
- Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal SummarizationLitian Zhang, Xiaoming Zhang, Junshu PanAAAI 2022 · 53 citations
- Nice Perfume. How Long Did You Marinate in It? Multimodal Sarcasm ExplanationPoorav Desai, Tanmoy Chakraborty, Md. Shad AkhtarAAAI 2022 · 49 citations
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 46 citations
- Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationYunlong Liang, Fandong Meng, Jinan Xu, Jiaan Wang et al.ACL 2023 · 17 citations
Builds on1
Related papers
- Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 VideosNayu Liu, Kaiwen Wei, Xian Sun, Hongfeng Yu et al.EMNLP 2022 · 10 citations
- Denoising Bottleneck with Mutual Information Maximization for Video Multimodal FusionShaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu et al.ACL 2023 · 11 citations
- Multimodal Video Summarization via Time-Aware TransformersXindi Shang, Zehuan Yuan, Anran Wang, Changhu WangACM MM 2021 · 31 citations
- Cross-modal Fusion Transformer for Integrating Retrieved Knowledge into Video Caption GenerationKarina Abubakirova, Waseem Ullah, Latif U. Khan, Mohsen GuizaniKDD 2026
- A Knowledge Augmented and Multimodal-Based Framework for Video SummarizationJiehang Xie, Xuanbai Chen, Shao-Ping Lu, Yulu YangACM MM 2022 · 11 citations
