MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
Yingying Wang, Xuanhua He, Chen Wu, Jialing Huang, Suiyun Zhang, Rui Liu, Xinghao Ding, Haoxuan Che
Abstract
Pan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the two modalities. Traditional CNN-based methods typically rely on channel-wise concatenation with fixed convolutional operators, which limits their adaptability to diverse spatial and spectral variations. While cross-attention mechanisms enable global interactions, they are computationally inefficient and may dilute fine-grained correspondences, making it difficult to capture complex semantic relationships. Recent advances in the Multimodal Diffusion Transformer (MMDiT) architecture have demonstrated impressive success in image generation and editing tasks. Unlike cross-attention, MMDiT employs in-context conditioning to facilitate more direct and efficient cross-modal information exchange. In this paper, we propose MMMamba, a cross-modal in-context fusion framework for pan-sharpening, with the flexibility to support image super-resolution in a zero-shot manner. Built upon the Mamba architecture, our design ensures linear computational complexity while maintaining strong cross-modal interaction capacity. Furthermore, we introduce a novel multimodal interleaved (MI) scanning mechanism that facilitates effective information exchange between the PAN and MS modalities. Extensive experiments demonstrate the superior performance of our method compared to existing state-of-the-art (SOTA) techniques across multiple tasks and benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ef05901-084b-4a2d-98ee-202f04b7abfcBuilds on15
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- HorNet: Efficient High-Order Spatial Interactions with Recursive Gated ConvolutionsYongming Rao, Wenliang Zhao, Yansong Tang, Jie Zhou et al.NeurIPS 2022 · 422 citations
- Pan-Sharpening with Customized Transformer and Invertible Neural NetworkMan Zhou, Jie Huang, Yanchi Fang, Xueyang Fu et al.AAAI 2022 · 130 citations
Related papers
- CTCP: Cross Transformer and CNN for PansharpeningZhao Su, Yong Yang, Shuying Huang, Weiguo Wan et al.ACM MM 2023 · 7 citations
- MFmamba: A Multi-function Network for Panchromatic Image Resolution Restoration Based on State-Space ModelQian Jiang, Qianqian Wang, Xin Jin, Michal Wozniak et al.AAAI 2026 · 2 citations
- Accelerated Diffusion via High-Low Frequency Decomposition for Pan-SharpeningGe Meng, Jingjia Huang, Jingyan Tu, Yingying Wang et al.AAAI 2025 · 3 citations
- CrosST: Cross Swin 4D Transformer for Multi-Modal Alzheimer's DetectionHao Wang, Hanxiao Li, Li XuACM MM 2025 · 1 citation
- A Novel State Space Model with Local Enhancement and State Sharing for Image FusionZihan Cao, Xiao Wu, Liang-Jian Deng, Yu ZhongACM MM 2024 · 24 citations
