Capsule-based Object Tracking with Natural Language Specification
Ding Ma, Xiangqian Wu
Abstract
Tracking with Natural-Language Specification (TNL) is a joint topic of understanding the vision and natural language with a wide range of applications. In previous works, the communication between two heterogeneous features of vision and language is mainly through a simple dynamic convolution. However, the performance of prior works is capped by the difficulty of linguistic variation of natural language in modeling the dynamically changing target and its surroundings. In the meanwhile, natural language and vision are firstly fused and then utilized for tracking, which is hard to model the query-focused context. Query-focused should pay more attention to context modeling to promote the correlation between these two features. To address these issues, we propose a capsule-based network, referred to as CapsuleTNL, which performs regression tracking with natural language query. In the beginning, the visual and textual input is encoded with capsules, which can not only establish the relationship between entities but also the relationship between the parts of the entity itself. Then, we devise two interaction routing modules, which consist of visual-textual routing module to reduce the linguistic variation of input query and textual-visual routing module to precisely incorporate query-based visual cues simultaneously. To validate the potential of the proposed network for visual object tracking, we evaluate our method on two large tracking benchmarks. The experimental evaluation demonstrates the effectiveness of our capsule-based network.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a1c181aa-05dc-412e-b81c-abd2f60ff694Cited by top-tier papers8
- CiteTracker: Correlating Image and Text for Visual TrackingXin Li, Yuqing Huang, Zhenyu He, Yaowei Wang et al.ICCV 2023 · 75 citations
- All in One: Exploring Unified Vision-Language Tracking with Multi-Modal AlignmentChunhui Zhang, Xin Sun, Yiqian Yang, Li Liu et al.ACM MM 2023 · 41 citations
- HERO: HiErarchical spatio-tempoRal reasOning with Contrastive Action Correspondence for End-to-End Video Object GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACM MM 2022 · 25 citations
- Tracking by Natural Language Specification with Long Short-term Context DecouplingDing Ma, Xiangqian WuICCV 2023 · 22 citations
- SUTrack: Towards Simple and Unified Single Object TrackingXin Chen, Ben Kang, Wanting Geng, Jiawen Zhu et al.AAAI 2025 · 12 citations
Related papers
- Context-Aware Integration of Language and Visual References for Natural Language TrackingYanyan Shao, Shuting He, Qi Ye, Yuchao Feng et al.CVPR 2024 · 42 citations
- Visual-Textual Capsule Routing for Text-Based Video SegmentationBruce McIntosh, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahCVPR 2020
- Joint Visual Grounding and Tracking with Natural Language SpecificationLi Zhou, Zikun Zhou, Kaige Mao, Zhenyu HeCVPR 2023
- Linguistically Routing Capsule Network for Out-of-distribution Visual Question AnsweringQingxing Cao, Wentao Wan, Keze Wang, Xiaodan Liang et al.ICCV 2021 · 16 citations
- CapsuleRRT: Relationships-Aware Regression Tracking via CapsulesDing Ma, Xiangqian WuCVPR 2021
