AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
Yan Li, Yifei Xing, Xiangyuan Lan, Xin Li, Haifeng Chen, Dongmei Jiang
摘要
Cross-modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-based methods have shown promising results in modeling inter-modal relationships, their quadratic computational complexity limits their applicability to long-sequence or large-scale data. Although recent Mamba-based approaches achieve linear complexity, their sequential scanning mechanism poses fundamental challenges in comprehensively modeling cross-modal relationships. To address this limitation, we propose Align-Mamba, an efficient and effective method for multimodal fusion. Specifically, grounded in Optimal Transport, we introduce a local cross-modal alignment module that explicitly learns token-level correspondences between different modalities. Moreover, we propose a global cross-modal alignment loss based on Maximum Mean Discrepancy to implicitly enforce the consistency between different modal distributions. Finally, the unimodal representations after local and global alignment are passed to the Mamba backbone for further cross-modal interaction and multimodal fusion. Extensive experiments on complete and incomplete multimodal fusion tasks demonstrate the effectiveness and efficiency of the proposed method. For instance, on the CMU-MOSI dataset, AlignMamba improves classification accuracy by 0.9%, reduces GPU memory usage by 20.3%, and decreases inference time by 83.3%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Enhance-then-Balance Modality Collaboration for Robust Multimodal Sentiment AnalysisKang He, Yuzhe Ding, Xinrong Wang, Fei Li 等CVPR 2026 · 被引用 1 次
- LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural CracksHui Liu, Chen Jia, Fan Shi, Xu Cheng 等ACM MM 2025 · 被引用 1 次
- Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMsHongcheng Liu, Yuhao Wang, Zhe Chen, Pingjie Wang 等ACL 2026
- DuetMerging: Synergizing Dynamic and Static Strategies for Mitigating Task Interference in Model MergingYan Li, Guiping Cao, Yaguang Song, Ming Tao 等CVPR 2026
- Walking Further: Semantic-Aware Multimodal Gait Recognition Under Long-Range ConditionsZhiyang Lu, Wen Jiang, Tianren Wu, Zhichao Wang 等AAAI 2026
它引用的顶会 Paper19
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
相关 Paper
- EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical AlignmentYifei Xing, Xiangyuan Lan, Ruiping Wang, Dongmei Jiang 等ICLR 2025
- MSAmba: Exploring Multimodal Sentiment Analysis with State Space ModelsXilin He, Haijian Liang, Boyi Peng, Weicheng Xie 等AAAI 2025 · 被引用 14 次
- VAMBA: Understanding Hour-Long Videos with Hybrid Mamba-TransformersWeiming Ren, Wentao Ma, Huan Yang, Cong Wei 等ICCV 2025 · 被引用 2 次
- Self-supervised Multiplex Consensus Mamba for General Image FusionYingying Wang, Rongjin Zhuang, Hui Zheng, Xuanhua He 等AAAI 2026 · 被引用 2 次
- Coupled Mamba: Enhanced Multimodal Fusion with Coupled State Space ModelWenbing Li, Hang Zhou, Junqing Yu, Zikai Song 等NeurIPS 2024 · 被引用 71 次
