Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object Detection
Jinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu, Risheng Liu, Wei Zhong, Zhongxuan Luo
Abstract
This study addresses the issue of fusing infrared and visible images that appear differently for object detection. Aiming at generating an image of high visual quality, previous approaches discover commons underlying the two modalities and fuse upon the common space either by iterative optimization or deep networks. These approaches neglect that modality differences implying the complementary information are extremely important for both fusion and subsequent detection task. This paper proposes a bilevel optimization formulation for the joint problem of fusion and detection, and then unrolls to a target-aware Dual Adversarial Learning (TarDAL) network for fusion and a commonly used detection network. The fusion network with one generator and dual discriminators seeks commons while learning from differences, which preserves structural information of targets from the infrared and textural details from the visible. Furthermore, we build a synchronized imaging system with calibrated infrared and optical sensors, and collect currently the most comprehensive benchmark covering a wide range of scenarios. Extensive experiments on several public datasets and our benchmark demonstrate that our method outputs not only visually appealing fusion but also higher detection mAP than the state-of-the-art approaches. The source code and benchmark are available at https://github.com/dlut-dimt/TarDAL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers108
- DDFM: Denoising Diffusion Model for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang et al.ICCV 2023 · 350 citations
- Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and SegmentationJinyuan Liu, Zhu Liu, Guanyao Wu, Long Ma et al.ICCV 2023 · 287 citations
- Equivariant Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Jiangshe Zhang, Yulun Zhang et al.CVPR 2024 · 155 citations
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang et al.CVPR 2024 · 121 citations
- Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionBing Cao, Yiming Sun, Pengfei Zhu, Qinghua HuICCV 2023 · 110 citations
Builds on3
- Searching a Hierarchically Aggregated Fusion Architecture for Fast Multi-Modality Image FusionRisheng Liu, Zhu Liu, Jinyuan Liu, Xin FanACM MM 2021 · 64 citations
- Retinex-Inspired Unrolling With Cooperative Prior Architecture Search for Low-Light Image EnhancementRisheng Liu, Long Ma, Jiaao Zhang, Xin Fan et al.CVPR 2021
- Learning a Neural Solver for Multiple Object TrackingGuillem Brasó, Laura Leal-TaixéCVPR 2020
Related papers
- DetFusion: A Detection-driven Infrared and Visible Image Fusion NetworkYiming Sun, Bing Cao, Pengfei Zhu, Qinghua HuACM MM 2022 · 165 citations
- MetaFusion: Infrared and Visible Image Fusion via Meta-Feature Embedding from Object DetectionWenda Zhao, Shigeng Xie, Fan Zhao, You He et al.CVPR 2023
- DDFD: Diffusion-Based Denoising Fusion for Object Detection in Infrared-Visible ImagesMin Dang, Gang Liu, Jingqi Zhao, Adams Wai-Kin Kong et al.ACM MM 2025 · 3 citations
- Dispel Darkness for Better Fusion: A Controllable Visual Enhancer Based on Cross-Modal Conditional Adversarial LearningHao Zhang, Linfeng Tang, Xinyu Xiang, Xuhui Zuo et al.CVPR 2024 · 21 citations
- Infrared-Privileged UAV Detection via Cross-Modal Vector-QuantizationZhibo Lou, Ruijie Zhang, Zeyu Luo, Qianxi Cao et al.AAAI 2026
