Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
Xiantao Hu, Ying Tai, Xu Zhao, Chen Zhao, Zhenyu Zhang, Jun Li, Bineng Zhong, Jian Yang
摘要
Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal relationships between video frames. These approaches do not fully exploit the temporal correlations in multimodal videos, making it difficult to capture the dynamic changes and motion information of targets in complex scenarios. To alleviate this problem, we propose a unified multimodal spatial-temporal tracking approach named STTrack. In contrast to previous paradigms that solely relied on updating reference information, we introduced a temporal state generator (TSG) that continuously generates a sequence of tokens containing multimodal temporal information. These temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the target. Furthermore, at the spatial level, we introduced the mamba fusion and background suppression interactive (BSI) modules. These modules establish a dual-stage mechanism for coordinating information interaction and fusion between modalities. Extensive comparisons on five benchmark datasets illustrate that STTrack achieves state-of-the-art performance across various multimodal tracking scenarios. Code is available at: https://github.com/NJU-PCALab/STTrack .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality DatasetChen Zhao, En Ci, Yunzhe Xu, Tiehan Fan 等NeurIPS 2025 · 被引用 24 次
- Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer NetworkXiang Fang, Wanlong Fang, Changshuo Wang, Daizong Liu 等AAAI 2025 · 被引用 10 次
- What You Have is What You Track: Adaptive and Robust Multimodal TrackingYuedong Tan, Jiawei Shao, Eduard Zamfir, Ruanjun Li 等ICCV 2025 · 被引用 5 次
- Hypergraph-State Collaborative Reasoning for Multi-Object TrackingZikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 等CVPR 2026 · 被引用 4 次
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao 等AAAI 2026 · 被引用 4 次
它引用的顶会 Paper22
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 被引用 746 次
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu 等NeurIPS 2022 · 被引用 556 次
- Attribute-Based Progressive Fusion Network for RGBT TrackingYun Xiao, Mengmeng Yang, Chenglong Li, Lei Liu 等AAAI 2022 · 被引用 218 次
- Prompting for Multi-Modal TrackingJinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis 等ACM MM 2022 · 被引用 167 次
相关 Paper
- Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object TrackingShilei Wang, Pujian Lai, Dong Gao, Jifeng Ning 等AAAI 2026
- MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language TrackingXinqi Liu, Li Zhou, Zikun Zhou, Jianqiu Chen 等CVPR 2025
- Exploiting All Mamba Fusion for Efficient RGB-D TrackingGe Ying, Dawei Zhang, Chengzhuan Yang, Wei Liu 等AAAI 2026
- Exploring Enhanced Contextual Information for Video-Level Object TrackingBen Kang, Xin Chen, Simiao Lai, Yang Liu 等AAAI 2025 · 被引用 48 次
- Temporal Adaptive RGBT Tracking with Modality PromptHongyu Wang, Xiaotao Liu, Yifan Li, Meng Sun 等AAAI 2024 · 被引用 92 次
