CO²IF: Language-Bridging Hyperspectral-Multispectral Image Fusion with Coordinated and Cross-modal Optimal Transport
Mingjin Zhang, Zhongkai Yang, Fei Gao
Abstract
Due to the difficulties of directly obtaining high-resolution hyperspectral images (HR-HSI), the fusion of low-resolution hyperspectral images (LR-HSI) and high-resolution multispectral images (HR-MSI) has emerged as an effective approach. While existing methods leverage image-level priors from HR-MSI, they often lack explicit semantic guidance for precise detail reconstruction. Recognizing that textual scene descriptions encapsulate valuable object attributes and contextual information, we introduce the first Language-Bridging framework for Hyperspectral and Multispectral image fusion (CO 2 IF). CO 2 IF leverages language semantics as prior knowledge to explicitly guide the reconstruction process. To bridge the modality gap between textual descriptions and high-dimensional hyperspectral data, we design a Crossmodal Optimal Transport (COT) module. COT establishes precise semantic correspondences between language features and the visual cues of individual spectral bands. Building upon this semantic alignment, we develop a Multimodal Coordinated State Space Model (CoMamba). CoMamba effectively integrates the language-derived priors with spatial information from HR-MSI and spectral information from LR-HSI. This language-guided reconstruction significantly enhances the extraction of crucial spatial-spectral details, leading to superior fidelity in the generated HR-HSI. In addition, this paper adds text descriptions for three widely used datasets. Both qualitative and quantitative experimental results on the public datasets confirm the superiority of the proposed method compared to the SOTA methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d40e09e1-eba7-4471-822f-cfeade1a246bBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient InferenceHan Zhao, Min Zhang, Wei Zhao, Pengxiang Ding et al.AAAI 2025 · 125 citations
- HSR-Diff: Hyperspectral Image Super-Resolution via Conditional Diffusion ModelsChanyue Wu, Dong Wang, Yunpeng Bai, Hanyu Mao et al.ICCV 2023 · 78 citations
Related papers
- Infrared and Visible Image Fusion with Language-Driven Loss in CLIP Embedding SpaceYuhao Wang, Lingjuan Miao, Zhiqiang Zhou, Lei Zhang et al.ACM MM 2025 · 18 citations
- A Novel State Space Model with Local Enhancement and State Sharing for Image FusionZihan Cao, Xiao Wu, Liang-Jian Deng, Yu ZhongACM MM 2024 · 24 citations
- Sp3ctralMamba: Physics-Driven Joint State Space Model for Hyperspectral Image ReconstructionGe Meng, Jingyan Tu, Jingjia Huang, Yunlong Lin et al.AAAI 2025 · 9 citations
- Image Fusion via Vision-Language ModelZixiang Zhao, Lilun Deng, Haowen Bai, Yukun Cui et al.ICML 2024 · 79 citations
- MSAmba: Exploring Multimodal Sentiment Analysis with State Space ModelsXilin He, Haijian Liang, Boyi Peng, Weicheng Xie et al.AAAI 2025 · 14 citations
