Temporal Adaptive RGBT Tracking with Modality Prompt
Hongyu Wang, Xiaotao Liu, Yifan Li, Meng Sun, Dian Yuan, Jing Liu
Abstract
RGBT tracking has been widely used in various fields such as robotics, surveillance processing, and autonomous driving. Existing RGBT trackers fully explore the spatial information between the template and the search region and locate the target based on the appearance matching results. However, these RGBT trackers have very limited exploitation of temporal information, either ignoring temporal information or exploiting it through online sampling and training. The former struggles to cope with the object state changes, while the latter neglects the correlation between spatial and temporal information. To alleviate these limitations, we propose a novel Temporal Adaptive RGBT Tracking framework, named as TATrack. TATrack has a spatio-temporal two-stream structure and captures temporal information by an online updated template, where the two-stream structure refers to the multi-modal feature extraction and cross-modal interaction for the initial template and the online update template respectively. TATrack contributes to comprehensively exploit spatio-temporal information and multi-modal information for target localization. In addition, we design a spatio-temporal interaction (STI) mechanism that bridges two branches and enables cross-modal interaction to span longer time scales. Extensive experiments on three popular RGBT tracking benchmarks show that our method achieves state-of-the-art performance, while running at real-time speed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6142dc6-ca16-4f8e-beb2-06b0be9fa7cbCited by top-tier papers13
- Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingXiantao Hu, Ying Tai, Xu Zhao, Chen Zhao et al.AAAI 2025 · 65 citations
- Cross-modulated Attention Transformer for RGBT TrackingYun Xiao, Jiacong Zhao, Andong Lu, Chenglong Li et al.AAAI 2025 · 28 citations
- RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion MambaAndong Lu, Wanyu Wang, Chenglong Li, Jin Tang et al.AAAI 2025 · 22 citations
- Breaking Modality Gap in RGBT Tracking: Coupled Knowledge DistillationAndong Lu, Jiacong Zhao, Chenglong Li, Yun Xiao et al.ACM MM 2024 · 15 citations
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao et al.AAAI 2026 · 4 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
Related papers
- Bridging Search Region Interaction with Template for RGB-T TrackingTianrui Hui, Zizheng Xun, Fengguang Peng, Junshi Huang et al.CVPR 2023
- AlignTrack: Top-Down Spatiotemporal Resolution Alignment for RGB-Event Visual TrackingChuanyu Sun, Jiqing Zhang, Yang Wang, Yuanchen Wang et al.AAAI 2026
- CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal FeaturesXiaokun Feng, Dailing Zhang, Shiyu Hu, Xuchen Li et al.ICML 2025
- Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingLiqiu Chen, Yuqing Huang, Hengyu Li, Zikun Zhou et al.ACM MM 2024 · 2 citations
- RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented GenerationHao Li, Yuhao Wang, Wenning Hao, Pingping Zhang et al.CVPR 2026 · 2 citations
