ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
Xiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang, Meiqi Wu, Jing Zhang, Xiaotang Chen, Kaiqi Huang
Abstract
Vision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, it is essential not only to characterize the target features but also to utilize the context features related to the target. However, the visual and textual target-context cues derived from the initial prompts generally align only with the initial target state. Due to their dynamic nature, target states are constantly changing, particularly in complex long-term sequences. It is intractable for these cues to continuously guide Vision-Language Trackers (VLTs). Furthermore, for the text prompts with diverse expressions, our experiments reveal that existing VLTs struggle to discern which words pertain to the target or the context, complicating the utilization of textual cues. In this work, we present a novel tracker named ATCTrack, which can obtain multimodal cues Aligned with the dynamic target states through comprehensive Target-Context feature modeling, thereby achieving robust tracking. Specifically, (1) for the visual modality, we propose an effective temporal visual target-context modeling approach that provides the tracker with timely visual cues. (2) For the textual modality, we achieve precise target words identification solely based on textual content, and design an innovative context words calibration method to adaptively utilize auxiliary context words. (3) We conduct extensive experiments on mainstream benchmarks and ATCTrack achieves a new SOTA performance. The code and models will be released at: https://github.com/XiaokunFeng/ATCTrack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b634033-eaaa-442f-a989-4a6b137897a0Cited by top-tier papers5
- NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video GenerationXiaokun Feng, Haiming Yu, Meiqi Wu, Shiyu Hu et al.ICLR 2026 · 13 citations
- CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT TrackingHao Li, Yuhao Wang, Xiantao Hu, Wenning Hao et al.AAAI 2026 · 4 citations
- RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented GenerationHao Li, Yuhao Wang, Wenning Hao, Pingping Zhang et al.CVPR 2026 · 2 citations
- MVLM: Template-Free Tracking via Vision-Language Margin Confidence and Memory-Gated TrackingDae-Hyeon Park, Mina Baek, Jeong-Hun Ha, Chan-Seop Park et al.CVPR 2026
- Interactive Tracking: A Human-in-the-Loop Paradigm with Memory-Augmented AdaptationYuqing Huang, Guotian Zeng, Zhenqiao Yuan, Zhenyu He et al.CVPR 2026
Builds on41
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Learning Target Candidate Association to Keep Track of What Not to TrackChristoph Mayer, Martin Danelljan, Danda Pani Paudel, Luc Van GoolICCV 2021 · 356 citations
Related papers
- Dynamic Updates for Language Adaptation in Visual-Language TrackingXiaohai Li, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.CVPR 2025
- MemVLT: Vision-Language Tracking with Adaptive Memory-based PromptsXiaokun Feng, Xuchen Li, Shiyu Hu, Dailing Zhang et al.NeurIPS 2024 · 34 citations
- Aware Distillation for Robust Vision-Language Tracking Under Linguistic SparsityGuangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang et al.AAAI 2026
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object TrackingKaiyang Lan, Ying Cui, Chenchen Jing, Jianwei Zheng et al.CVPR 2026
- ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelYiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang et al.NeurIPS 2024 · 21 citations
