FD2-Net: Frequency-Driven Feature Decomposition Network for Infrared-Visible Object Detection
Ke Li, Di Wang, Zhangyuan Hu, Shaofeng Li, Weiping Ni, Lin Zhao, Quan Wang
Abstract
Infrared-visible object detection (IVOD) seeks to harness the complementary information in infrared and visible images, thereby enhancing the performance of detectors in complex environments. However, existing methods often neglect the frequency characteristics of complementary information, such as the abundant high-frequency details in visible images and the valuable low-frequency thermal information in infrared images, thus constraining detection performance. To solve this problem, we introduce a novel Frequency-Driven Feature Decomposition Network for IVOD, called FD 2 -Net, which effectively captures the unique frequency representations of complementary information across multimodal visual spaces. Specifically, we propose a feature decomposition encoder, wherein the high-frequency unit (HFU) utilizes discrete cosine transform to capture representative highfrequency features, while the low-frequency unit (LFU) employs dynamic receptive fields to model the multi-scale context of diverse objects. Next, we adopt a parameter-free complementary strengths strategy to enhance multimodal features through seamless inter-frequency recoupling. Furthermore, we innovatively design a multimodal reconstruction mechanism that recovers image details lost during feature extraction, further leveraging the complementary information from infrared and visible images to enhance overall representational capacity. Extensive experiments demonstrate that FD 2 -Net outperforms state-of-the-art (SOTA) models across various IVOD benchmarks, i.e. LLVIP (96.2% mAP), FLIR (82.9% mAP), and M 3 FD (83.5% mAP).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba99d510-a5b1-4155-ae1a-9271bb851818Cited by top-tier papers5
- Rethinking Multi-Modal Object Detection From the Perspective of Mono-Modality Feature LearningTianyi Zhao, Boyang Liu, Yanglei Gao, Yiming Sun et al.ICCV 2025 · 15 citations
- Robust Pedestrian Detection with Uncertain ModalityQian Bie, Xiao Wang, Bin Yang, Zhixi Yu et al.AAAI 2026
- DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object DetectionChaolang Li, Pengwen Dai, Jingyu Li, Siyuan Yao et al.CVPR 2026
- ControlFuse: Instruction-guided Multi-Granularity Controllable Image FusionLibo Zhao, Xiaoli Zhang, Zeyu WangAAAI 2026
- UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV DetectionShenghui Huang, Menghao Hu, Longkun Zou, Hongyu Chi et al.CVPR 2026
Builds on12
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu et al.CVPR 2022 · 929 citations
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution BlocksXiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong HanICCV 2019 · 845 citations
- Large Selective Kernel Network for Remote Sensing Object DetectionYuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng et al.ICCV 2023 · 535 citations
- DDFM: Denoising Diffusion Model for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang et al.ICCV 2023 · 350 citations
- Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and SegmentationJinyuan Liu, Zhu Liu, Guanyao Wu, Long Ma et al.ICCV 2023 · 287 citations
Related papers
- Prior-Constrained Relevant Feature driven Image Fusion with Hybrid Feature via Mode DecompositionBingfeng Liu, Songwei Pei, Shuhuai Wang, Wenzheng Yang et al.ACM MM 2025
- DetFusion: A Detection-driven Infrared and Visible Image Fusion NetworkYiming Sun, Bing Cao, Pengfei Zhu, Qinghua HuACM MM 2022 · 165 citations
- DDFD: Diffusion-Based Denoising Fusion for Object Detection in Infrared-Visible ImagesMin Dang, Gang Liu, Jingqi Zhao, Adams Wai-Kin Kong et al.ACM MM 2025 · 3 citations
- WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object DetectionHaodong Zhu, Wenhao Dong, Linlin Yang, Hong Li et al.ICCV 2025 · 38 citations
- TIRDet: Mono-Modality Thermal InfraRed Object Detection Based on Prior Thermal-To-Visible TranslationZeyu Wang, Fabien Colonnier, Jinghong Zheng, Jyotibdha Acharya et al.ACM MM 2023 · 28 citations
