Aware Distillation for Robust Vision-Language Tracking Under Linguistic Sparsity
Guangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang, Tian Bai
Abstract
Vision-language object tracking overcomes the limitations of relying solely on visual features by leveraging language descriptions of objects to provide cross-modal semantic information, thereby enhancing model robustness in complex scenarios. However, most existing high-performance vision-language trackers are trained jointly on pure visual data and vision-language multimodal data. Due to the relative sparsity of language annotations in the data, the trackers tend to prioritize the localization role of visual features, diminishing the model's attention to language information. To mitigate this issue, we propose a novel vision-language tracker: Aware Distillation for Robust Vision-Language Tracking under Linguistic Sparsity (ADTrack). We introduce a knowledge distillation framework employing a knowledge-rich teacher model and a lightweight student model to establish modality correlations between vision and language, enabling efficient modeling between visual information and language descriptions. Specifically, our lightweight student module simultaneously distills language encoding capabilities from large language models through teacher-guided learning on input language, while performing target-aware perception on template images using language descriptions to generate more effective template features for subsequent visual extraction. Furthermore, to ensure perceptual robustness in linguistically sparse scenarios, we simulate language-deficient conditions during training and employ contrastive learning to enhance model adaptability. Extensive experiments demonstrate that ADTrack reduces parameters by over 50% while achieving state-of-the-art (SOTA) performance and speed on vision-language tracking benchmarks, including LaSOT, LaSOText, TNL2K, OTB-Lang and MGIT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d17ac14f-9b5c-419b-ba03-3bd8d31aa7fcBuilds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Divert More Attention to Vision-Language TrackingMingzhe Guo, Zhipeng Zhang, Heng Fan, Liping JingNeurIPS 2022 · 122 citations
Related papers
- Dynamic Updates for Language Adaptation in Visual-Language TrackingXiaohai Li, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.CVPR 2025
- All in One: Exploring Unified Vision-Language Tracking with Multi-Modal AlignmentChunhui Zhang, Xin Sun, Yiqian Yang, Li Liu et al.ACM MM 2023 · 41 citations
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingXiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang et al.ICCV 2025 · 3 citations
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object TrackingKaiyang Lan, Ying Cui, Chenchen Jing, Jianwei Zheng et al.CVPR 2026
- Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object TrackingGuangyao Li, Siping Zhuang, Yajun Jian, Yan Yan et al.ICCV 2025 · 8 citations
