TeRF: Text-driven and Region-aware Flexible Visible and Infrared Image Fusion
Hebaixu Wang, Hao Zhang, Xunpeng Yi, Xinyu Xiang, Leyuan Fang, Jiayi Ma
Abstract
The fusion of visible and infrared images aims to produce high-quality fusion images with rich textures and salient target information. Existing methods lack interactivity and flexibility in the execution of fusion. It is unfeasible to express the requirements to modify the fusion effect, and the different regions in the source images are treated equally across the identical fusion model, which causes fusion homogenization and low distinction. Besides, their pre-defined fusion strategies invariably lead to monotonous effects, which are insufficiently comprehensive. They fail to adequately consider data credibility, scene illumination, and noise degradation inherent in the source information. To address these issues, we propose the Te xt-driven and Region-aware Flexible visible and infrared image fusion, termed as TeRF. On the one hand, we propose a flexible image fusion framework with multiple large language and vision models, which facilitates the visual-text interaction. On the other hand, we aggregate comprehensive fine-tuning paradigms for the different fusion requirements to build a unified fine-tuning pipeline. It allows the linguistic selection of the regions and effects, yielding visually appealing fusion outcomes. Extensive experiments demonstrate the competitiveness of our method both qualitatively and quantitatively compared to existing state-of-the-art methods. Our code is publicly available at https://github.com/Baixuzx7/TeRF.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d6694eed-93ec-4f3f-b8cb-50f8908bbef9Cited by top-tier papers4
- LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up TablesXunpeng Yi, Yibing Zhang, Xinyu Xiang, Qinglong Yan et al.ICCV 2025 · 11 citations
- Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image FusionZeyu Wang, Jizheng Zhang, Haiyu Song, Mingyu Ge et al.ICCV 2025 · 10 citations
- LSFDNet: A Single-Stage Fusion and Detection Network for Ships Using SWIR and LWIRYanyin Guo, Runxuan An, Junwei Li, Zhiyuan ZhangACM MM 2025 · 2 citations
- MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven SemanticsJing Li, Yifan Wang, Jiafeng Yan, Renlong Zhang et al.AAAI 2026
Related papers
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang et al.CVPR 2024 · 121 citations
- Infrared and Visible Image Fusion with Language-Driven Loss in CLIP Embedding SpaceYuhao Wang, Lingjuan Miao, Zhiqiang Zhou, Lei Zhang et al.ACM MM 2025 · 18 citations
- Image Fusion via Vision-Language ModelZixiang Zhao, Lilun Deng, Haowen Bai, Yukun Cui et al.ICML 2024 · 79 citations
- ControlFuse: Instruction-guided Multi-Granularity Controllable Image FusionLibo Zhao, Xiaoli Zhang, Zeyu WangAAAI 2026
- Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionBing Cao, Yiming Sun, Pengfei Zhu, Qinghua HuICCV 2023 · 110 citations
