Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
Xunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang, Jiayi Ma
Abstract
Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve them, we introduce a novel approach that leverages semantic text guidance image fusion model for degradation-aware and interactive image fusion task, termed as Text-IF. It innovatively extends the classical image fusion to the text guided image fusion along with the ability to harmoniously address the degradation and interaction issues during fusion. Through the text semantic encoder and semantic interaction fusion decoder, Text-IF is accessible to the all-in-one infrared and visible image degradation-aware processing and the interactive flexible fusion outcomes. In this way, Text-IF achieves not only multi-modal image fusion, but also multi-modal information fusion. Extensive experiments prove that our proposed text guided image fusion strategy has obvious advantages over SOTA methods in the image fusion performance and degradation treatment. The code is available at https://github.com/XunpengYi/Text-IF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers49
- WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object DetectionHaodong Zhu, Wenhao Dong, Linlin Yang, Hong Li et al.ICCV 2025 · 38 citations
- Conditional Controllable Image FusionBing Cao, Xingxin Xu, Pengfei Zhu, Qilong Wang et al.NeurIPS 2024 · 29 citations
- Infrared and Visible Image Fusion with Language-Driven Loss in CLIP Embedding SpaceYuhao Wang, Lingjuan Miao, Zhiqiang Zhou, Lei Zhang et al.ACM MM 2025 · 18 citations
- Efficient Rectified Flow for Image FusionZirui Wang, Jiayi Zhang, Tianwei Guan, Yuhan Zhou et al.NeurIPS 2025 · 16 citations
- LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up TablesXunpeng Yi, Yibing Zhang, Xinyu Xiang, Qinglong Yan et al.ICCV 2025 · 11 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
Related papers
- Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion ModelHao Zhang, Lei Cao, Jiayi MaNeurIPS 2024 · 69 citations
- TeRF: Text-driven and Region-aware Flexible Visible and Infrared Image FusionHebaixu Wang, Hao Zhang, Xunpeng Yi, Xinyu Xiang et al.ACM MM 2024 · 11 citations
- MMAIF: Multi-Task and Multi-Degradation All-in-One for Image Fusion with Language GuidanceZihan Cao, Yu Zhong, Ziqi Wang, Liang-Jian DengICCV 2025 · 1 citation
- ControlFusion: A Controllable Image Fusion Network with Language-Vision Degradation PromptsLinfeng Tang, Yeda Wang, Zhanchuan Cai, Junjun Jiang et al.NeurIPS 2025 · 7 citations
- DRMF: Degradation-Robust Multi-Modal Image Fusion via Composable Diffusion PriorLinfeng Tang, Yuxin Deng, Xunpeng Yi, Qinglong Yan et al.ACM MM 2024 · 53 citations
