Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning
Yaozong Zheng, Qihua Liang, Bineng Zhong, Shuimu Zeng, Yuanliang Xue, Ning Li, Shuxiang Song
Abstract
Learning robust contextual knowledge from unlabeled videos is essential for advancing self-supervised tracking. However, conventional self-supervised trackers lack effective context modeling, while existing context association methods based on non-semantic queries struggle to adapt to unlabeled tracking scenarios, making it difficult to learn reliable contextual cues. In this work, we propose a novel self-supervised tracking framework, named ****, which introduces a dual-modal context association mechanism that jointly leverages fine-grained semantic prompts and contextual noise to drive the model toward learning robust tracking representations. Adherent to the easy-to-hard learning principle, our contextual association mechanism operates based on two stages. During early training, instance patch tokens (prompts) are assigned to both forward and backward tracking branches to facilitate the acquisition of tracking knowledge. As training progresses, contextual noise is gradually injected into the model to perturb feature, encouraging the tracker to learn robust tracking representations in a more complex feature space. Thus, this novel contextual association mechanism enables our self-supervised model to learn high-quality tracking representations from unlabeled videos, while being applied exclusively during training to preserve efficient inference. Extensive experiments demonstrate the superiority of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0ce26ac-7616-4df8-8533-ba5c3b18d440Builds on24
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Correlation-Aware Deep TrackingFei Xie, Chunyu Wang, Guangting Wang, Yue Cao et al.CVPR 2022 · 189 citations
- Autoregressive Queries for Adaptive Tracking with Spatio-Temporal TransformersJinxia Xie, Bineng Zhong, Zhiyi Mo, Shengping Zhang et al.CVPR 2024 · 100 citations
Related papers
- Unsupervised Learning of Accurate Siamese TrackingQiuhong Shen, Lei Qiao, Jinyang Guo, Peixia Li et al.CVPR 2022 · 73 citations
- S2SiamFC: Self-supervised Fully Convolutional Siamese Network for Visual TrackingChon-Hou Sio, Yu-Jen Ma, Hong-Han Shuai, Jun-Cheng Chen et al.ACM MM 2020 · 45 citations
- Exploiting Self-Supervised and Semi-Supervised Learning for Facial Landmark Tracking with Unlabeled DataShi Yin, Shangfei Wang, Xiaoping Chen, Enhong ChenACM MM 2020 · 7 citations
- Learning to Track Instance from Single Nature Language DescriptionYaozong Zheng, Bineng Zhong, Qihua Liang, Shuimu Zeng et al.CVPR 2026 · 1 citation
- Progressive Unsupervised Learning for Visual Object TrackingQiangqiang Wu, Jia Wan, Antoni B. ChanCVPR 2021
