Learning to Track Instance from Single Nature Language Description
Yaozong Zheng, Bineng Zhong, Qihua Liang, Shuimu Zeng, Haiying Xia, Shuxiang Song
Abstract
How to achieve vision-language (VL) tracking using natural language descriptions from a video sequence without relying on any bounding-box ground truth? In this work, we achieve this goal by tackling self-supervised VL tracking, which aims to evaluate tracking capabilities guided by natural language descriptions. We introduce ****, a novel self-supervised VL tracker that is capable of tracking any referred object by a language description. Unlike traditional methods that equally fuse all language and visual tokens, we propose an efficient Dynamic Token Aggregation Module, which treats each visual token unequally. The module consists of three main steps: i) Based on an anchor token, it selects multiple important target tokens from the template frame. ii) The selected target tokens are merged according to their attention scores and aggregated into the language tokens, thereby eliminating redundant visual token noise and enhancing semantic alignment. iii) Finally, the fused language tokens serve as guiding signals to extract potential target tokens from the search frame and propagate them to subsequent frames, enhancing temporal prompts and encouraging the tracker to autonomously learn instance tracking from unlabeled videos. This new modeling approach enables the effective self-supervised learning of language-guided tracking representations without the need for large-scale bounding box annotations. Extensive experiments on VL tracking benchmarks show that surpasses SOTA self-supervised methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Divert More Attention to Vision-Language TrackingMingzhe Guo, Zhipeng Zhang, Heng Fan, Liping JingNeurIPS 2022 · 122 citations
Related papers
- Consistencies are All You Need for Semi-supervised Vision-Language TrackingJiawei Ge, Jiuxin Cao, Xuelin Zhu, Xinyu Zhang et al.ACM MM 2024 · 22 citations
- ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelYiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang et al.NeurIPS 2024 · 21 citations
- All in One: Exploring Unified Vision-Language Tracking with Multi-Modal AlignmentChunhui Zhang, Xin Sun, Yiqian Yang, Li Liu et al.ACM MM 2023 · 41 citations
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object TrackingKaiyang Lan, Ying Cui, Chenchen Jing, Jianwei Zheng et al.CVPR 2026
- Unifying Visual and Vision-Language Tracking via Contrastive LearningYinchao Ma, Yuyang Tang, Wenfei Yang, Tianzhu Zhang et al.AAAI 2024 · 63 citations
