DetFusion: A Detection-driven Infrared and Visible Image Fusion Network
Yiming Sun, Bing Cao, Pengfei Zhu, Qinghua Hu
Abstract
Infrared and visible image fusion aims to utilize the complementary information between the two modalities to synthesize a new image containing richer information. Most existing works have focused on how to better fuse the pixel-level details from both modalities in terms of contrast and texture, yet ignoring the fact that the significance of image fusion is to better serve downstream tasks. For object detection tasks, object-related information in images is often more valuable than focusing on the pixel-level details of images alone. To fill this gap, we propose a detection-driven infrared and visible image fusion network, termed DetFusion, which utilizes object-related information learned in the object detection networks to guide multimodal image fusion. We cascade the image fusion network with the detection networks of both modalities and use the detection loss of the fused images to provide guidance on task-related information for the optimization of the image fusion network. Considering that the object locations provide a priori information for image fusion, we propose an object-aware content loss that motivates the fusion model to better learn the pixel-level information in infrared and visible images. Moreover, we design a shared attention module to motivate the fusion network to learn object-specific information from the object detection networks. Extensive experiments show that our DetFusion outperforms state-of-the-art methods in maintaining pixel intensity distribution and preserving texture details. More notably, the performance comparison with state-of-the-art image fusion methods in task-driven evaluation also demonstrates the superiority of the proposed method. Our code will be available: https://github.com/SunYM2020/DetFusion.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d12e19b3-b32e-47d0-901d-7c75bdf7e435Cited by top-tier papers29
- Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and SegmentationJinyuan Liu, Zhu Liu, Guanyao Wu, Long Ma et al.ICCV 2023 · 287 citations
- Equivariant Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Jiangshe Zhang, Yulun Zhang et al.CVPR 2024 · 155 citations
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang et al.CVPR 2024 · 121 citations
- Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionBing Cao, Yiming Sun, Pengfei Zhu, Qinghua HuICCV 2023 · 110 citations
- Learning a Graph Neural Network with Cross Modality Interaction for Image FusionJiawei Li, Jiansheng Chen, Jinyuan Liu, Huimin MaACM MM 2023 · 85 citations
Related papers
- MetaFusion: Infrared and Visible Image Fusion via Meta-Feature Embedding from Object DetectionWenda Zhao, Shigeng Xie, Fan Zhao, You He et al.CVPR 2023
- Task-driven Image Fusion with Learnable Fusion LossHaowen Bai, Jiangshe Zhang, Zixiang Zhao, Yichen Wu et al.CVPR 2025
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- DDFD: Diffusion-Based Denoising Fusion for Object Detection in Infrared-Visible ImagesMin Dang, Gang Liu, Jingqi Zhao, Adams Wai-Kin Kong et al.ACM MM 2025 · 3 citations
- Dispel Darkness for Better Fusion: A Controllable Visual Enhancer Based on Cross-Modal Conditional Adversarial LearningHao Zhang, Linfeng Tang, Xinyu Xiang, Xuhui Zuo et al.CVPR 2024 · 21 citations
