Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
Xiantao Hu, Ying Tai, Xu Zhao, Chen Zhao, Zhenyu Zhang, Jun Li, Bineng Zhong, Jian Yang
Abstract
Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal relationships between video frames. These approaches do not fully exploit the temporal correlations in multimodal videos, making it difficult to capture the dynamic changes and motion information of targets in complex scenarios. To alleviate this problem, we propose a unified multimodal spatial-temporal tracking approach named STTrack. In contrast to previous paradigms that solely relied on updating reference information, we introduced a temporal state generator (TSG) that continuously generates a sequence of tokens containing multimodal temporal information. These temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the target. Furthermore, at the spatial level, we introduced the mamba fusion and background suppression interactive (BSI) modules. These modules establish a dual-stage mechanism for coordinating information interaction and fusion between modalities. Extensive comparisons on five benchmark datasets illustrate that STTrack achieves state-of-the-art performance across various multimodal tracking scenarios. Code is available at: https://github.com/NJU-PCALab/STTrack .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0d137e7-6915-4991-a86e-dba0734a408eCited by top-tier papers38
- UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality DatasetChen Zhao, En Ci, Yunzhe Xu, Tiehan Fan et al.NeurIPS 2025 · 24 citations
- Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer NetworkXiang Fang, Wanlong Fang, Changshuo Wang, Daizong Liu et al.AAAI 2025 · 10 citations
- What You Have is What You Track: Adaptive and Robust Multimodal TrackingYuedong Tan, Jiawei Shao, Eduard Zamfir, Ruanjun Li et al.ICCV 2025 · 5 citations
- Hypergraph-State Collaborative Reasoning for Multi-Object TrackingZikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang et al.CVPR 2026 · 4 citations
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao et al.AAAI 2026 · 4 citations
Builds on22
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu et al.NeurIPS 2022 · 556 citations
- Attribute-Based Progressive Fusion Network for RGBT TrackingYun Xiao, Mengmeng Yang, Chenglong Li, Lei Liu et al.AAAI 2022 · 218 citations
- Prompting for Multi-Modal TrackingJinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis et al.ACM MM 2022 · 167 citations
Related papers
- Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object TrackingShilei Wang, Pujian Lai, Dong Gao, Jifeng Ning et al.AAAI 2026
- MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language TrackingXinqi Liu, Li Zhou, Zikun Zhou, Jianqiu Chen et al.CVPR 2025
- Exploiting All Mamba Fusion for Efficient RGB-D TrackingGe Ying, Dawei Zhang, Chengzhuan Yang, Wei Liu et al.AAAI 2026
- Exploring Enhanced Contextual Information for Video-Level Object TrackingBen Kang, Xin Chen, Simiao Lai, Yang Liu et al.AAAI 2025 · 48 citations
- Temporal Adaptive RGBT Tracking with Modality PromptHongyu Wang, Xiaotao Liu, Yifan Li, Meng Sun et al.AAAI 2024 · 92 citations
