Denoising Bottleneck with Mutual Information Maximization for Video Multimodal Fusion
Shaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu, Binghuai Lin, Yunbo Cao, Zhifang Sui
Abstract
Video multimodal fusion aims to integrate multimodal signals in videos, such as visual, audio and text, to make a complementary prediction with multiple modalities contents. However, unlike other image-text multimodal tasks, video has longer multimodal sequences with more redundancy and noise in both visual and audio modalities. Prior denoising methods like forget gate are coarse in the granularity of noise filtering. They often suppress the redundant and noisy information at the risk of losing critical information. Therefore, we propose a denoising bottleneck fusion (DBF) model for fine-grained video multimodal fusion. On the one hand, we employ a bottleneck mechanism to filter out noise and redundancy with a restrained receptive field. On the other hand, we use a mutual information maximization module to regulate the filter-out module to preserve key information within different modalities. Our DBF model achieves significant improvement over current state-of-the-art baselines on multiple benchmarks covering multimodal sentiment analysis and multimodal summarization tasks. It proves that our model can effectively capture salient features from noisy and redundant video, audio, and text inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dabf7e05-4629-477d-8717-0095245f58f2Cited by top-tier papers3
- MSAmba: Exploring Multimodal Sentiment Analysis with State Space ModelsXilin He, Haijian Liang, Boyi Peng, Weicheng Xie et al.AAAI 2025 · 14 citations
- DGLF: A Dual Graph-based Learning Framework for Multi-modal Sarcasm DetectionZhihong Zhu, Kefan Shen, Zhaorun Chen, Yunyan Zhang et al.EMNLP 2024 · 5 citations
- Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent DetectionZhanpeng Chen, Zhihong Zhu, Xianwei Zhuang, Zhiqi Huang et al.EMNLP 2024 · 4 citations
Builds on8
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language AnalysisZhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, Yingyu LiangAAAI 2020 · 419 citations
Related papers
- Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain VideosNayu Liu, Xian Sun, Hongfeng Yu, Wenkai Zhang et al.EMNLP 2020 · 54 citations
- InMu-Net: Advancing Multi-modal Intent Detection via Information Bottleneck and Multi-sensory ProcessingZhihong Zhu, Xuxin Cheng, Zhaorun Chen, Yuyan Chen et al.ACM MM 2024 · 11 citations
- CCAF: Coarse-to-fine Cross-Modal Alignment and Fusion for Multimodal Sentiment AnalysisXianbing Zhao, Shengzun Yang, Buzhou TangWWW 2026
- DiffuFuse: Diffusion-Driven Dual-Stream Fusion Framework for Multimodal Sentiment AnalysisXiongjian Lv, Yimin Wen, Hang YuACM MM 2025
- MDF: A Modality-Aware Disentanglement and Fusion Framework for Multimodal Sentiment AnalysisZhongquan Jian, Wenhan Lv, Yanhao Chen, Guanran Luo et al.AAAI 2026
