Image Fusion via Vision-Language Model
Zixiang Zhao, Lilun Deng, Haowen Bai, Yukun Cui, Zhipeng Zhang, Yulun Zhang, Haotong Qin, Dongdong Chen, Jiangshe Zhang, Peng Wang, Luc Van Gool
Abstract
Image fusion integrates essential information from multiple images into a single composite, enhancing structures, textures, and refining imperfections. Existing methods predominantly focus on pixel-level and semantic visual features for recognition, but often overlook the deeper text-level semantic information beyond vision. Therefore, we introduce a novel fusion paradigm named image Fusion via vIsion-Language Model (FILM), for the first time, utilizing explicit textual information from source images to guide the fusion process. Specifically, FILM generates semantic prompts from images and inputs them into ChatGPT for comprehensive textual descriptions. These descriptions are fused within the textual domain and guide the visual information fusion, enhancing feature extraction and contextual understanding, directed by textual semantic information via cross-attention. FILM has shown promising results in four image fusion tasks: infrared-visible, medical, multi-exposure, and multi-focus image fusion. We also propose a vision-language dataset containing ChatGPT-generated paragraph descriptions for the eight image fusion datasets across four fusion tasks, facilitating future research in vision-language model-based image fusion. Code and dataset are available at https://github. com/Zhaozixiang1228/IF-FILM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9cfd8788-6eb0-4c01-b546-1a2559abfff5Cited by top-tier papers24
- A Unified Solution to Video Fusion: From Multi-Frame Learning to BenchmarkingZixiang Zhao, Haowen Bai, Bingxin Ke, Yukun Cui et al.NeurIPS 2025 · 21 citations
- Infrared and Visible Image Fusion with Language-Driven Loss in CLIP Embedding SpaceYuhao Wang, Lingjuan Miao, Zhiqiang Zhou, Lei Zhang et al.ACM MM 2025 · 18 citations
- Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image FusionZeyu Wang, Jizheng Zhang, Haiyu Song, Mingyu Ge et al.ICCV 2025 · 10 citations
- Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition CuesChen Chen, Kangcheng Bin, Ting Hu, Jiahao Qi et al.ICCV 2025 · 8 citations
- Retinex-MEF: Retinex-Based Glare Effects Aware Unsupervised Multi-Exposure Image FusionHaowen Bai, Jiangshe Zhang, Zixiang Zhao, Lilun Deng et al.ICCV 2025 · 8 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang et al.CVPR 2024 · 121 citations
- TeRF: Text-driven and Region-aware Flexible Visible and Infrared Image FusionHebaixu Wang, Hao Zhang, Xunpeng Yi, Xinyu Xiang et al.ACM MM 2024 · 11 citations
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination MitigationZheng Qi, Chao Shang, Evangelia Spiliopoulou, Nikolaos PappasICML 2026
- IRGPT: Understanding Real-World Infrared Image with Bi-Cross-Modal Curriculum on Large-Scale BenchmarkZhe Cao, Jin Zhang, Ruiheng ZhangICCV 2025 · 2 citations
- Text-Image Conditioned 3D GenerationJiazhong Cen, Jiemin Fang, Sikuang Li, Guanjun Wu et al.CVPR 2026 · 1 citation
