Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model
Hao Zhang, Lei Cao, Jiayi Ma
Abstract
Existing multi-modal image fusion methods fail to address the compound degradations presented in source images, resulting in fusion images plagued by noise, color bias, improper exposure, etc. Additionally, these methods often overlook the specificity of foreground objects, weakening the salience of the objects of interest within the fused images. To address these challenges, this study proposes a novel interactive multi-modal image fusion framework based on the text-modulated diffusion model, called Text-DiFuse. First, this framework integrates feature-level information integration into the diffusion process, allowing adaptive degradation removal and multi-modal information fusion. This is the first attempt to deeply and explicitly embed information fusion within the diffusion process, effectively addressing compound degradation in image fusion. Second, by embedding the combination of the text and zero-shot location model into the diffusion fusion process, a text-controlled fusion re-modulation strategy is developed. This enables user-customized text control to improve fusion performance and highlight foreground objects in the fused images. Extensive experiments on diverse public datasets show that our Text-DiFuse achieves state-of-the-art fusion performance across various scenarios with complex degradation. Moreover, the semantic segmentation experiment validates the significant enhancement in semantic performance achieved by our text-controlled fusion re-modulation strategy. The code is publicly available at https://github.com/Leiii-Cao/Text-DiFuse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 532251af-ae5d-476f-8d76-9f392846f77cCited by top-tier papers17
- A Unified Solution to Video Fusion: From Multi-Frame Learning to BenchmarkingZixiang Zhao, Haowen Bai, Bingxin Ke, Yukun Cui et al.NeurIPS 2025 · 21 citations
- Efficient Rectified Flow for Image FusionZirui Wang, Jiayi Zhang, Tianwei Guan, Yuhan Zhou et al.NeurIPS 2025 · 16 citations
- ControlFusion: A Controllable Image Fusion Network with Language-Vision Degradation PromptsLinfeng Tang, Yeda Wang, Zhanchuan Cai, Junjun Jiang et al.NeurIPS 2025 · 7 citations
- Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive LawsLin Guo, Xiaoqing Luo, Wei Xie, Zhancheng Zhang et al.NeurIPS 2025 · 5 citations
- Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation ScenariosYu Shi, Yu Liu, Zhong-Cheng Wu, Juan Cheng et al.CVPR 2026 · 4 citations
Builds on20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Denoising Diffusion Restoration ModelsBahjat Kawar, Michael Elad, Stefano Ermon, Jiaming SongNeurIPS 2022 · 1,439 citations
Related papers
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang et al.CVPR 2024 · 121 citations
- ReCoFuse: Ultra-Robust Image Fusion via Restorative Multi-Modal Diffusion Reciprocal CouplingHao Zhang, Shuhan Yang, Linfeng Tang, Xunpeng Yi et al.CVPR 2026
- DRMF: Degradation-Robust Multi-Modal Image Fusion via Composable Diffusion PriorLinfeng Tang, Yuxin Deng, Xunpeng Yi, Qinglong Yan et al.ACM MM 2024 · 53 citations
- UniLDiff: Unlocking the Power of Diffusion Priors for All-in-One Image RestorationZihan Cheng, Liangtai Zhou, Dian Chen, Ni Tang et al.CVPR 2026 · 5 citations
- MMAIF: Multi-Task and Multi-Degradation All-in-One for Image Fusion with Language GuidanceZihan Cao, Yu Zhong, Ziqi Wang, Liang-Jian DengICCV 2025 · 1 citation
