Semantic Feature Purification for Adversarially-Aware RGB-T Tracking
Jiahao Wang, Fang Liu, Hao Wang, Shuo Li, Xinyi Wang, Puhua Chen
摘要
RGB-T tracking is increasingly deployed in safety-critical applications such as autonomous driving, surveillance, and rescue robotics, where tracking reliability is essential under adverse conditions. Although the fusion of RGB and thermal infrared (TIR) modalities offers improved robustness in low-light and occluded scenes, recent findings show that RGB-T trackers remain highly susceptible to subtle input perturbations, human-imperceptible modifications that exploit cross-modal inconsistencies to mislead tracking outputs. In real-world scenarios, such perturbations can arise from sensor spoofing, infrared camouflage, or physical-world attacks, posing serious risks to operational safety. To address this, we propose SFPT, a Semantic Feature Purification framework that enhances RGB-T tracking at the representation level. Rather than filtering corrupted inputs at the pixel level, SFPT introduces task-specific semantic anchors into the feature space to reinforce perturbation-invariant cues. These anchors are derived from descriptive language, interact with visual features to purify representations. To further suppress modality-specific interference, we design an Adaptive Perturbation-Guided Cross-Modal Fusion (APG-CMF) module, which leverages language and visual signals to estimate reliability and dynamically reweight cross-modal features, ensuring robust fusion under perturbation conditions. Extensive experiments under diverse perturbation conditions validate the effectiveness of our approach. Notably, SFPT maintains performance comparable to clean settings even when subjected to perturbations of strength 1 255 and 4 255 , demonstrating strong resilience to real-world interference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal FusionSijie Mai, Haifeng Hu, Songlong XingAAAI 2020 · 被引用 233 次
- DepthTrack: Unveiling the Power of RGBD TrackingSong Yan, Jinyu Yang, Jani Käpylä, Feng Zheng 等ICCV 2021 · 被引用 114 次
相关 Paper
- Quality-Aware RGBT Tracking via Supervised Reliability Learning and Weighted Residual GuidanceLei Liu, Chenglong Li, Yun Xiao, Jin TangACM MM 2023 · 被引用 36 次
- ACAttack: Adaptive Cross Attacking RGB-T Tracker via Multi-Modal Response DecouplingXinyu Xiang, Qinglong Yan, Hao Zhang, Jiayi MaCVPR 2025
- Cross-Modal Stealth: A Coarse-to-Fine Attack Framework for RGB-T TrackerXinyu Xiang, Qinglong Yan, Hao Zhang, Jianfeng Ding 等AAAI 2025 · 被引用 3 次
- FA3T: Feature-Aware Adversarial Attacks for Multi-modal TrackingJiahao Wang, Fang Liu, Licheng Jiao, Hao Wang 等ACM MM 2025
- RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented GenerationHao Li, Yuhao Wang, Wenning Hao, Pingping Zhang 等CVPR 2026 · 被引用 2 次
