ConsMSA: Semantic Distribution Consistency Learning for Multimodal Sentiment Analysis
Pan Wang, Lipeng Ke, Huajun Ying, Pritish Mohapatra, Rohan Sarkar, Suresh Lakhani, sankar venkataraman, Jingtong Hu
Abstract
Multimodal sentiment analysis (MSA) aims to predict human sentiments by integrating signals from different modalities such as text, video, and audio. However, raw multimodal sequences often suffer from semantic inconsistency--exhibiting redundancy or conflicts within and across modalities--which hinders robust understanding and increases computational cost. To this end, we introduce ConsMSA, which explicitly formalizes semantic distribution consistency across both - and -modality, providing a consistency-aware mechanism for robust and efficient multimodal sentiment prediction. Specifically, ConsMSA projects multimodal token features into a shared sentiment space to compute an Intra- and Inter-modality Consistency Score (). By coupling this score with predictive relevance, we formulate consistency-aware importance signals that are utilized: (i) as a consistency regularizer to align latent distributions during training, (ii) to derive semantic-aware weights for adaptive multimodal token reweighting, and (iii) as a practical criterion to prune redundant or conflicting tokens. Extensive experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS demonstrate that ConsMSA achieves state-of-the-art performance while remaining robust under aggressive token compression--retaining only 10% of tokens yields comparable accuracy. These results establish semantic distribution consistency as a promising foundation for synergizing predictive robustness with computational efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef7d0871-04a5-4dfd-bdcc-6655d1c1d19dBuilds on13
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of ModalityWenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu et al.ACL 2020 · 376 citations
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
Related papers
- Cross-modality Representation Interactive Learning for Multimodal Sentiment AnalysisJian Huang, Yanli Ji, Yang Yang, Heng Tao ShenACM MM 2023 · 17 citations
- Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment AnalysisHaoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu et al.EMNLP 2023 · 131 citations
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu et al.CVPR 2026 · 7 citations
- CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment AnalysisHaoyu Jiang, Xiaoliang Chen, Duoqian Miao, Xiaolin Qin et al.CVPR 2026
- ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment AnalysisJiuding Yang, Yakun Yu, Di Niu, Weidong Guo et al.ACL 2023 · 135 citations
