Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and Benchmark
Xiao Wang, Xiujun Shu, Zhipeng Zhang, Bo Jiang, Yaowei Wang, Yonghong Tian, Feng Wu
Abstract
Tracking by natural language specification is a new rising research topic that aims at locating the target object in the video sequence based on its language description. Compared with traditional bounding box (BBox) based tracking, this setting guides object tracking with high-level semantic information, addresses the ambiguity of BBox, and links local and global search organically together. Those benefits may bring more flexible, robust and accurate tracking performance in practical scenarios. However, existing natural language initialized trackers are developed and compared on benchmark datasets proposed for trackingby-BBox, which can't reflect the true power of trackingby-language. In this work, we propose a new benchmark specifically dedicated to the tracking-by-language, including a large scale dataset, strong and diverse baseline methods. Specifically, we collect 2k video sequences (contains a total of 1,244,340 frames, 663 words) and split 1300/700 for the train/testing respectively. We densely annotate one sentence in English and corresponding bounding boxes of the target object for each video. We also introduce two new challenges into TNL2K for the object tracking task, i.e., adversarial samples and modality switch. A strong baseline method based on an adaptive local-global-search scheme is proposed for future works to compare. We believe this benchmark will greatly boost related researches on natural language guided tracking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers74
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.AAAI 2024 · 247 citations
- Learn to Match: Automatic Matching Network Design for Visual TrackingZhipeng Zhang, Yihao Liu, Xiao Wang, Bing Li et al.ICCV 2021 · 224 citations
- MixFormerV2: Efficient Fully Transformer TrackingYutao Cui, Tianhui Song, Gangshan Wu, Limin WangNeurIPS 2023 · 193 citations
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park et al.CVPR 2026 · 144 citations
- Divert More Attention to Vision-Language TrackingMingzhe Guo, Zhipeng Zhang, Heng Fan, Liping JingNeurIPS 2022 · 122 citations
Builds on13
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan et al.AAAI 2020 · 944 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Learning the Model Update for Siamese TrackersLichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer, Martin Danelljan et al.ICCV 2019 · 371 citations
- GlobalTrack: A Simple and Strong Baseline for Long-Term TrackingLianghua Huang, Xin Zhao, Kaiqi HuangAAAI 2020 · 278 citations
Related papers
- Context-Aware Integration of Language and Visual References for Natural Language TrackingYanyan Shao, Shuting He, Qi Ye, Yuchao Feng et al.CVPR 2024 · 42 citations
- Joint Visual Grounding and Tracking with Natural Language SpecificationLi Zhou, Zikun Zhou, Kaige Mao, Zhenyu HeCVPR 2023
- Unifying Visual and Vision-Language Tracking via Contrastive LearningYinchao Ma, Yuyang Tang, Wenfei Yang, Tianzhu Zhang et al.AAAI 2024 · 63 citations
- Tracking by Natural Language Specification with Long Short-term Context DecouplingDing Ma, Xiangqian WuICCV 2023 · 22 citations
- MVLM: Template-Free Tracking via Vision-Language Margin Confidence and Memory-Gated TrackingDae-Hyeon Park, Mina Baek, Jeong-Hun Ha, Chan-Seop Park et al.CVPR 2026
