Pseudo Visible Feature Fine-Grained Fusion for Thermal Object Detection
Ting Li, Mao Ye, Tianwen Wu, Nianxin Li, Shuaifeng Li, Song Tang, Luping Ji
Abstract
Thermal object detection is a critical task in various fields, such as surveillance and autonomous driving. Current state-of-the-art (SOTA) models always leverage a prior Thermal-To-Visible (T2V) translation model to obtain visible spectrum information, followed by a cross-modality aggregation module to fuse information from both modalities. However, this fusion approach does not fully exploit the complementary visible spectrum information beneficial for thermal detection. To address this issue, we propose a novel cross-modal fusion method called Pseudo Visible Feature Fine-Grained Fusion (PFGF). Specifically, a graph is constructed with nodes generated from multi-level thermal features and pseudo-visual latent features produced by the T2V model. Each level of features corresponds to a subgraph. An Inter-Mamba block is proposed to perform cross-modality fusion between nodes at the lowest level; while a Cascade Knowledge Integration (CKI) strategy is used to fuse low-level fused information to highlevel subgraphs in a cascade manner. After several iterations of graph node updating, each subgraph outputs an aggregated feature to the detection head respectively. Unlike previous cross-modal fusion methods, our approach explicitly models high-level relationships between crossmodal data, effectively fusing different granularity information. Experimental results demonstrate that our method achieves SOTA detection performance. Code is available at https://github.com/liting1018/PFGF .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7f14803-176e-4154-a204-d887e6aa49e6Cited by top-tier papers2
- DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain GuidanceYinghui Xing, Xiaoting Su, Shizhou Zhang, Donghao Chu et al.AAAI 2026
- No Calibration, No Depth, No Problem: Cross-Sensor View Synthesis with 3D ConsistencyCho-Ying Wu, Zixun Huang, Xinyu Huang, Liu RenCVPR 2026
Builds on12
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Learning a Graph Neural Network with Cross Modality Interaction for Image FusionJiawei Li, Jiansheng Chen, Jinyuan Liu, Huimin MaACM MM 2023 · 85 citations
- Source-Free Object Detection by Learning to Overlook Domain StyleShuaifeng Li, Mao Ye, Xiatian Zhu, Lihua Zhou et al.CVPR 2022 · 75 citations
- Robust Small-scale Pedestrian Detection with Cued Recall via Memory LearningJung Uk Kim, Sungjune Park, Yong Man RoICCV 2021 · 61 citations
Related papers
- TIRDet: Mono-Modality Thermal InfraRed Object Detection Based on Prior Thermal-To-Visible TranslationZeyu Wang, Fabien Colonnier, Jinghong Zheng, Jyotibdha Acharya et al.ACM MM 2023 · 28 citations
- DetFusion: A Detection-driven Infrared and Visible Image Fusion NetworkYiming Sun, Bing Cao, Pengfei Zhu, Qinghua HuACM MM 2022 · 165 citations
- SAM-Guided Semantic Knowledge Fusion for Visible-Infrared Object DetectionTing Li, Songtao Li, Shuaifeng Li, Xiaolin Qin et al.ACM MM 2025 · 2 citations
- MetaFusion: Infrared and Visible Image Fusion via Meta-Feature Embedding from Object DetectionWenda Zhao, Shigeng Xie, Fan Zhao, You He et al.CVPR 2023
- Dispel Darkness for Better Fusion: A Controllable Visual Enhancer Based on Cross-Modal Conditional Adversarial LearningHao Zhang, Linfeng Tang, Xinyu Xiang, Xuhui Zuo et al.CVPR 2024 · 21 citations
