MVLM: Template-Free Tracking via Vision-Language Margin Confidence and Memory-Gated Tracking
Dae-Hyeon Park, Mina Baek, Jeong-Hun Ha, Chan-Seop Park, Jamshidjon Ganiev, Seung-Hwan Bae
Abstract
We introduce a new template-free tracking paradigm based solely on natural language, capable of tracking an arbitrary object and seamlessly switching to a new target without box initialization. Our key idea is to localize an object via vision-language (VL) correlation. However, using the correlation alone is brittle under large search regions due to spatial uncertainty and ambiguous VL saliency. To resolve these, we propose MVLM, a memory-based vision-language margin confidence that integrates vision-language correlation, encoder prediction, and temporal memory. MVLM dynamically gates the search region-switching between compact ROI (Region of Interest) search and global relocalization-to reduce spatial uncertainty. Theoretically, we derive bounds that connect the MVLM score to tracking probability, characterizing mis-localization within ROI and ROI-exclusion probabilities. Through extensive evaluation, we validate our theorems and achieve state-of-the-art performance on several benchmarks (TNL2K, LaSOT, OTB99, and MGIT) using only language guidance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 280f9fc4-5ed0-4567-b837-3bf468b3c321Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Exploring Lightweight Hierarchical Vision Transformers for Efficient Visual TrackingBen Kang, Xin Chen, Dong Wang, Houwen Peng et al.ICCV 2023 · 119 citations
- Autoregressive Queries for Adaptive Tracking with Spatio-Temporal TransformersJinxia Xie, Bineng Zhong, Zhiyi Mo, Shengping Zhang et al.CVPR 2024 · 100 citations
Related papers
- Aware Distillation for Robust Vision-Language Tracking Under Linguistic SparsityGuangtong Zhang, Bineng Zhong, Shirui Yang, Yang Wang et al.AAAI 2026
- Dynamic Updates for Language Adaptation in Visual-Language TrackingXiaohai Li, Bineng Zhong, Qihua Liang, Zhiyi Mo et al.CVPR 2025
- MemVLT: Vision-Language Tracking with Adaptive Memory-based PromptsXiaokun Feng, Xuchen Li, Shiyu Hu, Dailing Zhang et al.NeurIPS 2024 · 34 citations
- Learning to Track Instance from Single Nature Language DescriptionYaozong Zheng, Bineng Zhong, Qihua Liang, Shuimu Zeng et al.CVPR 2026 · 1 citation
- Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and BenchmarkXiao Wang, Xiujun Shu, Zhipeng Zhang, Bo Jiang et al.CVPR 2021
