Tracking by Natural Language Specification with Long Short-term Context Decoupling
Ding Ma, Xiangqian Wu
摘要
The main challenge of Tracking by Natural Language Specification (TNL) is to predict the movement of the target object by giving two heterogeneous information, e.g., one is the static description of the main characteristics of a video contained in the textual query, i.e., long-term context; the other one is an image patch containing the object and its surroundings cropped from the current frame, i.e., the search area. Currently, most methods still struggle with the rationality of using those two information and simply fusing the two. However, the linguistic information contained in the textual query and the visual representation stored in the search area may sometimes be inconsistent, in which case the direct fusion of the two may lead to conflicts. To address this problem, we propose DecoupleTNL, introducing a video clip containing short-term context information into the framework of TNL and exploring a proper way to reduce the impact when visual representation is inconsistent with linguistic information. Concretely, we design two jointly optimized tasks, i.e., short-term context-matching and long-term context-perceiving. The context-matching task aims to gather the dynamic short-term context information in a period, while the context-perceiving task tends to extract the static long-term context information. After that, we design a long short-term modulation module to integrate both context information for accurate tracking. Extensive experiments have been conducted on three tracking benchmark datasets to demonstrate the superiority of DecoupleTNL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MemVLT: Vision-Language Tracking with Adaptive Memory-based PromptsXiaokun Feng, Xuchen Li, Shiyu Hu, Dailing Zhang 等NeurIPS 2024 · 被引用 34 次
- ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelYiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang 等NeurIPS 2024 · 被引用 21 次
- SUTrack: Towards Simple and Unified Single Object TrackingXin Chen, Ben Kang, Wanting Geng, Jiawen Zhu 等AAAI 2025 · 被引用 12 次
- ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingXiaokun Feng, Shiyu Hu, Xuchen Li, Dailing Zhang 等ICCV 2025 · 被引用 3 次
- UETrack: A Unified and Efficient Framework for Single Object TrackingBen Kang, Jie Zhao, Xin Chen, Wanting Geng 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper16
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 被引用 746 次
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani 等NeurIPS 2020 · 被引用 483 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
相关 Paper
- Context-Aware Integration of Language and Visual References for Natural Language TrackingYanyan Shao, Shuting He, Qi Ye, Yuchao Feng 等CVPR 2024 · 被引用 42 次
- Capsule-based Object Tracking with Natural Language SpecificationDing Ma, Xiangqian WuACM MM 2021 · 被引用 25 次
- Joint Visual Grounding and Tracking with Natural Language SpecificationLi Zhou, Zikun Zhou, Kaige Mao, Zhenyu HeCVPR 2023
- Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and BenchmarkXiao Wang, Xiujun Shu, Zhipeng Zhang, Bo Jiang 等CVPR 2021
- Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object TrackingKaiyang Lan, Ying Cui, Chenchen Jing, Jianwei Zheng 等CVPR 2026
