FD2-Net: Frequency-Driven Feature Decomposition Network for Infrared-Visible Object Detection
Ke Li, Di Wang, Zhangyuan Hu, Shaofeng Li, Weiping Ni, Lin Zhao, Quan Wang
摘要
Infrared-visible object detection (IVOD) seeks to harness the complementary information in infrared and visible images, thereby enhancing the performance of detectors in complex environments. However, existing methods often neglect the frequency characteristics of complementary information, such as the abundant high-frequency details in visible images and the valuable low-frequency thermal information in infrared images, thus constraining detection performance. To solve this problem, we introduce a novel Frequency-Driven Feature Decomposition Network for IVOD, called FD 2 -Net, which effectively captures the unique frequency representations of complementary information across multimodal visual spaces. Specifically, we propose a feature decomposition encoder, wherein the high-frequency unit (HFU) utilizes discrete cosine transform to capture representative highfrequency features, while the low-frequency unit (LFU) employs dynamic receptive fields to model the multi-scale context of diverse objects. Next, we adopt a parameter-free complementary strengths strategy to enhance multimodal features through seamless inter-frequency recoupling. Furthermore, we innovatively design a multimodal reconstruction mechanism that recovers image details lost during feature extraction, further leveraging the complementary information from infrared and visible images to enhance overall representational capacity. Extensive experiments demonstrate that FD 2 -Net outperforms state-of-the-art (SOTA) models across various IVOD benchmarks, i.e. LLVIP (96.2% mAP), FLIR (82.9% mAP), and M 3 FD (83.5% mAP).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Rethinking Multi-Modal Object Detection From the Perspective of Mono-Modality Feature LearningTianyi Zhao, Boyang Liu, Yanglei Gao, Yiming Sun 等ICCV 2025 · 被引用 15 次
- Robust Pedestrian Detection with Uncertain ModalityQian Bie, Xiao Wang, Bin Yang, Zhixi Yu 等AAAI 2026
- DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object DetectionChaolang Li, Pengwen Dai, Jingyu Li, Siyuan Yao 等CVPR 2026
- ControlFuse: Instruction-guided Multi-Granularity Controllable Image FusionLibo Zhao, Xiaoli Zhang, Zeyu WangAAAI 2026
- UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV DetectionShenghui Huang, Menghao Hu, Longkun Zou, Hongyu Chi 等CVPR 2026
它引用的顶会 Paper12
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu 等CVPR 2022 · 被引用 929 次
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution BlocksXiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong HanICCV 2019 · 被引用 845 次
- Large Selective Kernel Network for Remote Sensing Object DetectionYuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng 等ICCV 2023 · 被引用 535 次
- DDFM: Denoising Diffusion Model for Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang 等ICCV 2023 · 被引用 350 次
- Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and SegmentationJinyuan Liu, Zhu Liu, Guanyao Wu, Long Ma 等ICCV 2023 · 被引用 287 次
相关 Paper
- Prior-Constrained Relevant Feature driven Image Fusion with Hybrid Feature via Mode DecompositionBingfeng Liu, Songwei Pei, Shuhuai Wang, Wenzheng Yang 等ACM MM 2025
- DetFusion: A Detection-driven Infrared and Visible Image Fusion NetworkYiming Sun, Bing Cao, Pengfei Zhu, Qinghua HuACM MM 2022 · 被引用 165 次
- DDFD: Diffusion-Based Denoising Fusion for Object Detection in Infrared-Visible ImagesMin Dang, Gang Liu, Jingqi Zhao, Adams Wai-Kin Kong 等ACM MM 2025 · 被引用 3 次
- WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object DetectionHaodong Zhu, Wenhao Dong, Linlin Yang, Hong Li 等ICCV 2025 · 被引用 38 次
- TIRDet: Mono-Modality Thermal InfraRed Object Detection Based on Prior Thermal-To-Visible TranslationZeyu Wang, Fabien Colonnier, Jinghong Zheng, Jyotibdha Acharya 等ACM MM 2023 · 被引用 28 次
