ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model
Yiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang, Junwei Huang, Yang Li, Chenhui Li, Changbo Wang
Abstract
Visual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language (VL) trackers have proposed to utilize additional natural language descriptions to enhance versatility in various applications. However, VL trackers are still inferior to State-of-The-Art (SoTA) visual trackers in terms of tracking performance. We found that this inferiority primarily results from their heavy reliance on manual textual annotations, which include the frequent provision of ambiguous language descriptions. In this paper, we propose ChatTracker to leverage the wealth of world knowledge in the Multimodal Large Language Model (MLLM) to generate high-quality language descriptions and enhance tracking performance. To this end, we propose a novel reflection-based prompt optimization module to iteratively refine the ambiguous and inaccurate descriptions of the target with tracking feedback. To further utilize semantic information produced by MLLM, a simple yet effective VL tracking framework is proposed and can be easily integrated as a plug-and-play module to boost the performance of both VL and visual trackers. Experimental results show that our proposed ChatTracker achieves a performance comparable to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention CollapseYiming Sun, Mi Zhang, Feifei Li, Geng Hong et al.AAAI 2026 · 5 citations
- FARTrack: Fast Autoregressive Visual Tracking with High PerformanceGuijie Wang, Tong Lin, Yifan Bai, Anjia Cao et al.ICLR 2026 · 3 citations
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingXiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang et al.ICCV 2025 · 3 citations
- RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented GenerationHao Li, Yuhao Wang, Wenning Hao, Pingping Zhang et al.CVPR 2026 · 2 citations
- Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMsYikang Zhou, Tao Zhang, Shilin Xu, Shihao Chen et al.ICCV 2025 · 2 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Dynamic Updates for Language Adaptation in Visual-Language TrackingXiaohai Li, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.CVPR 2025
- Learning to Track Instance from Single Nature Language DescriptionYaozong Zheng, Bineng Zhong, Qihua Liang, Shuimu Zeng et al.CVPR 2026 · 1 citation
- Aware Distillation for Robust Vision-Language Tracking Under Linguistic SparsityGuangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang et al.AAAI 2026
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object TrackingKaiyang Lan, Ying Cui, Chenchen Jing, Jianwei Zheng et al.CVPR 2026
- Unifying Visual and Vision-Language Tracking via Contrastive LearningYinchao Ma, Yuyang Tang, Wenfei Yang, Tianzhu Zhang et al.AAAI 2024 · 63 citations
