MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
Yingying Wang, Xuanhua He, Chen Wu, Jialing Huang, Suiyun Zhang, Rui Liu, Xinghao Ding, Haoxuan Che
摘要
Pan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the two modalities. Traditional CNN-based methods typically rely on channel-wise concatenation with fixed convolutional operators, which limits their adaptability to diverse spatial and spectral variations. While cross-attention mechanisms enable global interactions, they are computationally inefficient and may dilute fine-grained correspondences, making it difficult to capture complex semantic relationships. Recent advances in the Multimodal Diffusion Transformer (MMDiT) architecture have demonstrated impressive success in image generation and editing tasks. Unlike cross-attention, MMDiT employs in-context conditioning to facilitate more direct and efficient cross-modal information exchange. In this paper, we propose MMMamba, a cross-modal in-context fusion framework for pan-sharpening, with the flexibility to support image super-resolution in a zero-shot manner. Built upon the Mamba architecture, our design ensures linear computational complexity while maintaining strong cross-modal interaction capacity. Furthermore, we introduce a novel multimodal interleaved (MI) scanning mechanism that facilitates effective information exchange between the PAN and MS modalities. Extensive experiments demonstrate the superior performance of our method compared to existing state-of-the-art (SOTA) techniques across multiple tasks and benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- HorNet: Efficient High-Order Spatial Interactions with Recursive Gated ConvolutionsYongming Rao, Wenliang Zhao, Yansong Tang, Jie Zhou 等NeurIPS 2022 · 被引用 422 次
- Pan-Sharpening with Customized Transformer and Invertible Neural NetworkMan Zhou, Jie Huang, Yanchi Fang, Xueyang Fu 等AAAI 2022 · 被引用 130 次
相关 Paper
- CTCP: Cross Transformer and CNN for PansharpeningZhao Su, Yong Yang, Shuying Huang, Weiguo Wan 等ACM MM 2023 · 被引用 7 次
- MFmamba: A Multi-function Network for Panchromatic Image Resolution Restoration Based on State-Space ModelQian Jiang, Qianqian Wang, Xin Jin, Michal Wozniak 等AAAI 2026 · 被引用 2 次
- Accelerated Diffusion via High-Low Frequency Decomposition for Pan-SharpeningGe Meng, Jingjia Huang, Jingyan Tu, Yingying Wang 等AAAI 2025 · 被引用 3 次
- CrosST: Cross Swin 4D Transformer for Multi-Modal Alzheimer's DetectionHao Wang, Hanxiao Li, Li XuACM MM 2025 · 被引用 1 次
- A Novel State Space Model with Local Enhancement and State Sharing for Image FusionZihan Cao, Xiao Wu, Liang-Jian Deng, Yu ZhongACM MM 2024 · 被引用 24 次
