CiteTracker: Correlating Image and Text for Visual Tracking
Xin Li, Yuqing Huang, Zhenyu He, Yaowei Wang, Huchuan Lu, Ming-Hsuan Yang
Abstract
Existing visual tracking methods typically take an image patch as the reference of the target to perform tracking. However, a single image patch cannot provide a complete and precise concept of the target object as images are limited in their ability to abstract and can be ambiguous, which makes it difficult to track targets with drastic variations. In this paper, we propose the CiteTracker to enhance target modeling and inference in visual tracking by connecting images and text. Specifically, we develop a text generation module to convert the target image patch into a descriptive text containing its class and attribute information, providing a comprehensive reference point for the target. In addition, a dynamic description module is designed to adapt to target variations for more effective target representation. We then associate the target description and the search image using an attention-based correlation module to generate the correlated features for target state reference. Extensive experiments on five diverse datasets are conducted to evaluate the proposed algorithm and the favorable performance against the state-of-the-art methods demonstrates the effectiveness of the proposed tracking method. The source code and trained models will be released at https: //github.com/NorahGreen/CiteTracker .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3163e64a-1c4e-4694-92d9-7daf50a8bba1Cited by top-tier papers19
- Exploring Enhanced Contextual Information for Video-Level Object TrackingBen Kang, Xin Chen, Simiao Lai, Yang Liu et al.AAAI 2025 · 48 citations
- Context-Aware Integration of Language and Visual References for Natural Language TrackingYanyan Shao, Shuting He, Qi Ye, Yuchao Feng et al.CVPR 2024 · 42 citations
- MemVLT: Vision-Language Tracking with Adaptive Memory-based PromptsXiaokun Feng, Xuchen Li, Shiyu Hu, Dailing Zhang et al.NeurIPS 2024 · 34 citations
- ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelYiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang et al.NeurIPS 2024 · 21 citations
- SUTrack: Towards Simple and Unified Single Object TrackingXin Chen, Ben Kang, Wanting Geng, Jiawen Zhu et al.AAAI 2025 · 12 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
Related papers
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingXiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang et al.ICCV 2025 · 3 citations
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object TrackingKaiyang Lan, Ying Cui, Chenchen Jing, Jianwei Zheng et al.CVPR 2026
- Dynamic Updates for Language Adaptation in Visual-Language TrackingXiaohai Li, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.CVPR 2025
- Generalizing Multiple Object Tracking to Unseen Domains by Introducing Natural Language RepresentationEn Yu, Songtao Liu, Zhuoling Li, Jinrong Yang et al.AAAI 2023 · 22 citations
- Learning to Track Instance from Single Nature Language DescriptionYaozong Zheng, Bineng Zhong, Qihua Liang, Shuimu Zeng et al.CVPR 2026 · 1 citation
