Bridging Search Region Interaction with Template for RGB-T Tracking
Tianrui Hui, Zizheng Xun, Fengguang Peng, Junshi Huang, Xiaoming Wei, Xiaolin Wei, Jiao Dai, Jizhong Han, Si Liu
Abstract
RGB-T tracking aims to leverage the mutual enhancement and complement ability of RGB and TIR modalities for improving the tracking process in various scenarios, where cross-modal interaction is the key component. Some previous methods concatenate the RGB and TIR search region features directly to perform a coarse interaction process with redundant background noises introduced. Many other methods sample candidate boxes from search frames and conduct various fusion approaches on isolated pairs of RGB and TIR boxes, which limits the cross-modal interaction within local regions and brings about inadequate context modeling. To alleviate these limitations, we propose a novel Template-Bridged Search region Interaction (TBSI) module which exploits templates as the medium to bridge the crossmodal interaction between RGB and TIR search regions by gathering and distributing target-relevant object and environment contexts. Original templates are also updated with enriched multimodal contexts from the template medium. Our TBSI module is inserted into a ViT backbone for joint feature extraction, search-template matching, and crossmodal interaction. Extensive experiments on three popular RGB-T tracking benchmarks demonstrate our method achieves new state-of-the-art performances. Code is available at https://github.com/RyanHTR/TBSI .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a353f82c-edec-4762-802a-f92fa4c2fba5Cited by top-tier papers26
- Temporal Adaptive RGBT Tracking with Modality PromptHongyu Wang, Xiaotao Liu, Yifan Li, Meng Sun et al.AAAI 2024 · 92 citations
- Generative-Based Fusion Mechanism for Multi-Modal TrackingZhangyong Tang, Tianyang Xu, Xiaojun Wu, Xuefeng Zhu et al.AAAI 2024 · 78 citations
- Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingXiantao Hu, Ying Tai, Xu Zhao, Chen Zhao et al.AAAI 2025 · 65 citations
- Cross-modulated Attention Transformer for RGBT TrackingYun Xiao, Jiacong Zhao, Andong Lu, Chenglong Li et al.AAAI 2025 · 28 citations
- RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion MambaAndong Lu, Wanyu Wang, Chenglong Li, Jin Tang et al.AAAI 2025 · 22 citations
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan et al.AAAI 2020 · 944 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
Related papers
- Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingLiqiu Chen, Yuqing Huang, Hengyu Li, Zikun Zhou et al.ACM MM 2024 · 2 citations
- Cross-Modal Pattern-Propagation for RGB-T TrackingChaoqun Wang, Chunyan Xu, Zhen Cui, Ling Zhou et al.CVPR 2020
- Cross-modal Orthogonal High-rank Augmentation for RGB-Event Transformer-trackersZhiyu Zhu, Junhui Hou, Dapeng Oliver WuICCV 2023 · 66 citations
- Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel ApproachYun Xiao, Yuhang Wang, Jiandong Jin, Wankang Zhang et al.AAAI 2026
- ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic SegmentationQiang Zhang, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang et al.CVPR 2021
