Context-Aware Integration of Language and Visual References for Natural Language Tracking
Yanyan Shao, Shuting He, Qi Ye, Yuchao Feng, Wenhan Luo, Jiming Chen
摘要
Tracking by natural language specification (TNL) aims to consistently localize a target in a video sequence given a linguistic description in the initial frame. Existing methodologies perform language-based and template-based matching for target reasoning separately and merge the matching results from two sources, which suffer from tracking drift when language and visual templates missalign with the dynamic target state and ambiguity in the later merging stage. To tackle the issues, we propose a joint multi-modal tracking framework with 1) a prompt modulation module to leverage the complementarity between temporal visual templates and language expressions, enabling precise and context-aware appearance and linguistic cues, and 2) a unified target decoding module to integrate the multi-modal reference cues and executes the integrated queries on the search image to predict the target location in an end-to-end manner directly. This design ensures spatio-temporal consistency by leveraging historical visual information and introduces an integrated solution, generating predictions in a single step. Extensive experiments conducted on TNL2K, OTB-Lang, LaSOT, and RefCOCOg validate the efficacy of our proposed approach. The results demonstrate competitive performance against state-of-the-art methods for both tracking and grounding. Code is available at https://github.com/twotw02/QueryNLT
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- MemVLT: Vision-Language Tracking with Adaptive Memory-based PromptsXiaokun Feng, Xuchen Li, Shiyu Hu, Dailing Zhang 等NeurIPS 2024 · 被引用 34 次
- RefMask3D: Language-Guided Transformer for 3D Referring SegmentationShuting He, Henghui DingACM MM 2024 · 被引用 12 次
- Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object TrackingGuangyao Li, Siping Zhuang, Yajun Jian, Yan Yan 等ICCV 2025 · 被引用 8 次
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingXiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang 等ICCV 2025 · 被引用 3 次
- Learning to Track Instance from Single Nature Language DescriptionYaozong Zheng, Bineng Zhong, Qihua Liang, Shuimu Zeng 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- SwinTrack: A Simple and Strong Baseline for Transformer TrackingLiting Lin, Heng Fan, Zhipeng Zhang, Yong Xu 等NeurIPS 2022 · 被引用 556 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
相关 Paper
- Joint Visual Grounding and Tracking with Natural Language SpecificationLi Zhou, Zikun Zhou, Kaige Mao, Zhenyu HeCVPR 2023
- Tracking by Natural Language Specification with Long Short-term Context DecouplingDing Ma, Xiangqian WuICCV 2023 · 被引用 22 次
- Capsule-based Object Tracking with Natural Language SpecificationDing Ma, Xiangqian WuACM MM 2021 · 被引用 25 次
- Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and BenchmarkXiao Wang, Xiujun Shu, Zhipeng Zhang, Bo Jiang 等CVPR 2021
- Unifying Visual and Vision-Language Tracking via Contrastive LearningYinchao Ma, Yuyang Tang, Wenfei Yang, Tianzhu Zhang 等AAAI 2024 · 被引用 63 次
