Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment Analysis
Wei Han, Hui Chen, Soujanya Poria
Abstract
In multimodal sentiment analysis (MSA), the performance of a model highly depends on the quality of synthesized embeddings. These embeddings are generated from the upstream process called multimodal fusion, which aims to extract and combine the input unimodal raw data to produce a richer multimodal representation. Previous work either back-propagates the task loss or manipulates the geometric property of feature spaces to produce favorable fusion results, which neglects the preservation of critical task-related information that flows from input to the fusion results. In this work, we propose a framework named MultiModal InfoMax (MMIM), which hierarchically maximizes the Mutual Information (MI) in unimodal input pairs (inter-modality) and between multimodal fusion result and unimodal input in order to maintain taskrelated information through multimodal fusion. The framework is jointly trained with the main task (MSA) to improve the performance of the downstream MSA task. To address the intractable issue of MI bounds, we further formulate a set of computationally simple parametric and non-parametric methods to approximate their truth value. Experimental results on the two widely used datasets demonstrate the efficacy of our approach. The implementation of this work is publicly available at https://github.com/ declare-lab/Multimodal-Infomax.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers70
- UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion RecognitionGuimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu et al.EMNLP 2022 · 206 citations
- Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment AnalysisHaoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu et al.EMNLP 2023 · 131 citations
- FuseMoE: Mixture-of-Experts Transformers for Fleximodal FusionXing Han, Huy Nguyen, Carl Harris, Nhat Ho et al.NeurIPS 2024 · 129 citations
- Towards Robust Multimodal Sentiment Analysis with Incomplete DataHaoyu Zhang, Wenbin Wang, Tianshu YuNeurIPS 2024 · 90 citations
- Revisiting Disentanglement and Fusion on Modality and Context in Conversational Multimodal Emotion RecognitionBobo Li, Hao Fei, Lizi Liao, Yu Zhao et al.ACM MM 2023 · 76 citations
Builds on8
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language AnalysisZhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, Yingyu LiangAAAI 2020 · 419 citations
Related papers
- Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation LearningMingcheng Li, Dingkang Yang, Yang Liu, Shunli Wang et al.NeurIPS 2024 · 48 citations
- KEBR: Knowledge Enhanced Self-Supervised Balanced Representation for Multimodal Sentiment AnalysisAoqiang Zhu, Min Hu, Xiaohua Wang, Jiaoyun Yang et al.ACM MM 2024 · 12 citations
- Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment AnalysisZhilu Yang, Mingcheng LiCVPR 2026
- PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment AnalysisHeng Xie, Kang Zhu, Zhengqi Wen, Jianhua Tao et al.AAAI 2026 · 1 citation
- Multimodal fusion via cortical network inspired lossesShiv ShankarACL 2022 · 21 citations
