Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image Fusion
Zeyu Wang, Jizheng Zhang, Haiyu Song, Mingyu Ge, Jiayu Wang, Haoran Duan
Abstract
Infrared and visible image fusion (VIS-IR) aims to integrate complementary information from both source images to produce a fused image with enriched details. However, most existing fusion models lack controllability, making it difficult to customize the fused output according to user preferences. To address this challenge, we propose a novel weakly-supervised, instance-level controllable fusion model that adaptively highlights user-specified instances based on input text. Our model consists of two stages: pseudolabel generation and fusion network training. In the first stage, guided by observed multimodal manifold priors, we leverage text and manifold similarity as joint supervisory signals to train text-to-image response network (TIRN) in a weakly-supervised manner, enabling it to identify referenced semantic-level objects from instance segmentation outputs. To align text and image features in TIRN, we propose a multimodal feature alignment module (MFA), using manifold similarity to guide attention weight assignment for precise correspondence between image patches and text embeddings. Moreover, we employ spatial positional relationships to accurately select the referenced instances from multiple semantic-level objects. In the second stage, the fusion network takes source images and text as input, using the generated pseudo-labels for supervision to apply distinct fusion strategies for target and non-target regions. Experimental results show that our model achieves state-of-the-art fusion performance and accurately highlights user-defined instances. Code: https://github.com/GMY628/RIS-Fuse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c861cfd-256e-4d93-9ee4-5f2121d9036bCited by top-tier papers5
- HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-ResolutionYang Zou, Xingyue Zhu, Kaiqi Han, Jun Ma et al.AAAI 2026 · 3 citations
- ControlFuse: Instruction-guided Multi-Granularity Controllable Image FusionLibo Zhao, Xiaoli Zhang, Zeyu WangAAAI 2026
- SigFusion: Unified Signal-Level Self-Supervised Learning Paradigm for Image FusionZeyu Wang, Jiawei Feng, Jiayu Wang, Pengjie Wang et al.AAAI 2026
- Domain Adaptation Guided Infrared and Visible Image FusionTianwei Guan, Haozhen Wei, Yuhan Zhou, Jun Ma et al.AAAI 2026
- Breaking Task Boundaries: A Unified Model for 3D Medical Image Fusion and Segmentation Guided by Manifold PerspectiveZeyu Wang, Jiayu Wang, Haiyu SongAAAI 2026
Builds on18
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 457 citations
- DDFM: Denoising Diffusion Model for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang et al.ICCV 2023 · 350 citations
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao et al.CVPR 2022 · 337 citations
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
Related papers
- CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image FusionYiming Sun, Yuan Ruan, Qinghua Hu, Pengfei ZhuAAAI 2026
- TeRF: Text-driven and Region-aware Flexible Visible and Infrared Image FusionHebaixu Wang, Hao Zhang, Xunpeng Yi, Xinyu Xiang et al.ACM MM 2024 · 11 citations
- Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic SpaceHuan Kang, Hui Li, Tianyang Xu, Tao Zhou et al.ICML 2026
- DetFusion: A Detection-driven Infrared and Visible Image Fusion NetworkYiming Sun, Bing Cao, Pengfei Zhu, Qinghua HuACM MM 2022 · 165 citations
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang et al.CVPR 2024 · 121 citations
