Cross-modulated Attention Transformer for RGBT Tracking
Yun Xiao, Jiacong Zhao, Andong Lu, Chenglong Li, Bing Yin, Yin Lin, Cong Liu
摘要
Existing Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-modal feature interaction and template-search correlation computation. Nevertheless, the independent searchtemplate correlation calculations ignore the consistency between branches, which can result in ambiguous and inappropriate correlation weights. It not only limits the intramodal feature representation, but also harms the robustness of cross-attention for multi-modal feature interaction and search-template correlation computation. To address these issues, we propose a novel approach called Crossmodulated Attention Transformer (CAFormer), which performs intra-modality self-correlation, inter-modality feature interaction, and search-template correlation computation in a unified attention model, for RGBT tracking. In particular, we first independently generate correlation maps for each modality and feed them into the designed Correlation Modulated Enhancement module, modulating inaccurate correlation weights by seeking the consensus between modalities. Such kind of design unifies self-attention and cross-attention schemes, which not only alleviates inaccurate attention weight computation in self-attention but also eliminates redundant computation introduced by extra cross-attention scheme. In addition, we propose a collaborative token elimination strategy to further improve tracking inference efficiency and accuracy. Extensive experiments on five public RGBT tracking benchmarks show the outstanding performance of the proposed CAFormer against stateof-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao 等AAAI 2026 · 被引用 4 次
- SEATrack: Simple, Efficient, and Adaptive Multimodal TrackerJunbin Su, Ziteng Xue, Shihui Zhang, Kun Chen 等CVPR 2026 · 被引用 3 次
- Adaptive Depth Lightweight RGB-T Tracking with Holistic Token RoutingTian Ding, Hongtao Yang, Liangtao Shi, Jun Li 等CVPR 2026 · 被引用 3 次
- RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented GenerationHao Li, Yuhao Wang, Wenning Hao, Pingping Zhang 等CVPR 2026 · 被引用 2 次
- Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT TrackingAndong Lu, Ziyi Zha, Jiandong Jin, Shihao Li 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 被引用 746 次
- Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New BaselinePengyu Zhang, Jie Zhao, Dong Wang, Huchuan Lu 等CVPR 2022 · 被引用 225 次
相关 Paper
- Transformer TrackingXin Chen, Bin Yan, Jiawen Zhu, Dong Wang 等CVPR 2021
- Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingLiqiu Chen, Yuqing Huang, Hengyu Li, Zikun Zhou 等ACM MM 2024 · 被引用 2 次
- Temporal Adaptive RGBT Tracking with Modality PromptHongyu Wang, Xiaotao Liu, Yifan Li, Meng Sun 等AAAI 2024 · 被引用 92 次
- Generalized Relation Modeling for Transformer TrackingShenyuan Gao, Chunluan Zhou, Jun ZhangCVPR 2023
- UTPTrack: Towards Simple and Unified Token Pruning for Visual TrackingHao Wu, Xudong Wang, Jialiang Zhang, Junlong Tong 等CVPR 2026 · 被引用 6 次
