MSAmba: Exploring Multimodal Sentiment Analysis with State Space Models
Xilin He, Haijian Liang, Boyi Peng, Weicheng Xie, Muhammad Haris Khan, Siyang Song, Zitong Yu
Abstract
Multimodal sentiment analysis, which learns a model to process multiple modalities simultaneously and predict a sentiment value, is an important area of affective computing. Modeling sequential intra-modal information and enhancing cross-modal interactions are crucial to multimodal sentiment analysis. In this paper, we propose MSAmba, a novel hybrid Mamba-based architecture for multimodal sentiment analysis, consisting of two core blocks: Intra-Modal Sequential Mamba (ISM) block and Cross-Modal Hybrid Mamba (CHM) block, to comprehensively address the abovementioned challenges with hybrid state space models. Firstly, the ISM block models the sequential information within each modality in a bi-directional manner with the assistance of global information. Subsequently, the CHM blocks explicitly model centralized cross-modal interaction with a hybrid combination of Mamba and attention mechanism to facilitate information fusion across modalities. Finally, joint learning of the intra-modal tokens and cross-modal tokens is utilized to predict the sentiment values. This paper serves as one of the pioneering works to unravel the outstanding performances and great research potential of Mamba-based methods in the task of multimodal sentiment analysis. Experiments on CMU-MOSI, CMU-MOSEI and CH-SIMS demonstrate the superior performance of the proposed MSAmba over prior Transformer-based and CNN-based methods. Code is available at here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91f201a8-e499-464f-8512-182494a822fcCited by top-tier papers3
- QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment AnalysisYitong Zhu, Yuxuan Jiang, Guanxuan Jiang, Bojing Hou et al.ACL 2026
- Recovering Coherent Affective Patterns: Addressing Modality Missing in Multimodal Sentiment AnalysisHuiting Huang, Tieliang Gong, Kai He, Wen Wen et al.AAAI 2026
- From Static to Active: Knowledge-Aware Node State Selection in Multi-view Graph LearningWeiran Liao, Jielong Lu, Yuhong Chen, Shide Du et al.AAAI 2026
Builds on18
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
Related papers
- Coupled Mamba: Enhanced Multimodal Fusion with Coupled State Space ModelWenbing Li, Hang Zhou, Junqing Yu, Zikai Song et al.NeurIPS 2024 · 71 citations
- DDSE: A Decoupled Dual-Stream Enhanced Framework for Multimodal Sentiment Analysis with Text-Centric SSMShenjie Jiang, Zhuoyu Wang, Xuecheng Wu, Hongru Ji et al.ACM MM 2025 · 4 citations
- A Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment AnalysisDongning Rao, Yunbiao Zeng, Zhihua Jiang, Jujian LvAAAI 2026
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu et al.CVPR 2026 · 7 citations
- Dynamically Adjust Word Representations Using Unaligned Multimodal InformationJiwei Guo, Jiajia Tang, Weichen Dai, Yu Ding et al.ACM MM 2022 · 59 citations
