Coupled Mamba: Enhanced Multimodal Fusion with Coupled State Space Model
Wenbing Li, Hang Zhou, Junqing Yu, Zikai Song, Wei Yang
摘要
The essence of multi-modal fusion lies in exploiting the complementary information inherent in diverse modalities. However, prevalent fusion methods rely on traditional neural architectures and are inadequately equipped to capture the dynamics of interactions across modalities, particularly in presence of complex intra- and inter-modality correlations. Recent advancements in State Space Models (SSMs), notably exemplified by the Mamba model, have emerged as promising contenders. Particularly, its state evolving process implies stronger modality fusion paradigm, making multi-modal fusion on SSMs an appealing direction. However, fusing multiple modalities is challenging for SSMs due to its hardware-aware parallelism designs. To this end, this paper proposes the Coupled SSM model, for coupling state chains of multiple modalities while maintaining independence of intra-modality state processes. Specifically, in our coupled scheme, we devise an inter-modal hidden states transition scheme, in which the current state is dependent on the states of its own chain and that of the neighbouring chains at the previous time-step. To fully comply with the hardware-aware parallelism, we devise an expedite coupled state transition scheme and derive its corresponding global convolution kernel for parallelism. Extensive experiments on CMU-MOSEI, CH-SIMS, CH-SIMSV2 through multi-domain input verify the effectiveness of our model compared to current state-of-the-art methods, improved F1-Score by 0.4%, 0.9%, and 2.3% on the three datasets respectively, 49% faster inference and 83.7% GPU memory save. The results demonstrate that Coupled Mamba model is capable of enhanced multi-modal fusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Logical Phase Transitions: Understanding Collapse in LLM Logical ReasoningXinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu 等ACL 2026 · 被引用 28 次
- Semantic-Aware Logical Reasoning via a Semiotic FrameworkYunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li 等ACL 2026 · 被引用 27 次
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang 等AAAI 2026 · 被引用 24 次
- LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention RoutingWenbing Li, Zikai Song, Hang Zhou, Junqing Yu 等ICLR 2026 · 被引用 20 次
- ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Mingyu Zhang 等CVPR 2026 · 被引用 16 次
它引用的顶会 Paper22
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
相关 Paper
- MSAmba: Exploring Multimodal Sentiment Analysis with State Space ModelsXilin He, Haijian Liang, Boyi Peng, Weicheng Xie 等AAAI 2025 · 被引用 14 次
- PVMamba: Parallelizing Vision Mamba via Dynamic State AggregationFei Xie, Zhongdao Wang, Weijia Zhang, Chao MaICCV 2025 · 被引用 2 次
- Spatial-Mamba: Effective Visual State Space Models via Structure-Aware State FusionChaodong Xiao, Minghan Li, Zhengqiang Zhang, Deyu Meng 等ICLR 2025
- AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal AlignmentYan Li, Yifei Xing, Xiangyuan Lan, Xin Li 等CVPR 2025
- DDSE: A Decoupled Dual-Stream Enhanced Framework for Multimodal Sentiment Analysis with Text-Centric SSMShenjie Jiang, Zhuoyu Wang, Xuecheng Wu, Hongru Ji 等ACM MM 2025 · 被引用 4 次
