QuadTreeCapsule: QuadTree Capsules for Deep Regression Tracking
Ding Ma, Xiangqian Wu
Abstract
Benefit from the capability of capturing part-to-whole relationships, Capsule Network has been successful in many vision tasks. However, their high computational complexity poses a significant obstacle to applying them to visual tracking, requiring fast inference. In this paper, we introduce the idea of QuadTree Capsules, which explores the property of part-to-whole relationships endowed by the Capsule Network by significantly reducing the computational complexity. We build capsule pyramids and select meaningful relationships in a coarse-to-fine manner, dubbed as QuadTreeCapsule. Specifically, the top K capsules with the highest activation values are selected, and routing is only calculated within the relevant regions corresponding to these top K capsules with a novel symmetric guided routing algorithm. Additionally, considering the importance of temporal relationships, a multi-spectral pose matrix attention mechanism is developed for more accurate spatio-temporal capsule assignments between two sets of capsules. Moreover, during online inference, we shift part of the spatio-temporal capsules long the temporal dimension, facilitating information exchanged among neighboring frames. Extensive experimentation has proved the effectiveness of our methodology, which achieves state-of-the-art results compared with other tracking methods on eight widely-used benchmarks. Our tracker runs at approximately 43 fps on GPU.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ea086baf-8a1c-497c-a7ff-8461788f716dRelated papers
- CapsuleRRT: Relationships-Aware Regression Tracking via CapsulesDing Ma, Xiangqian WuCVPR 2021
- PT-CapsNet: A Novel Prediction-Tuning Capsule Network Suitable for Deeper ArchitecturesChenbin Pan, Senem VelipasalarICCV 2021 · 11 citations
- Quadtree Attention for Vision TransformersShitao Tang, Jiahui Zhang, Siyu Zhu, Ping TanICLR 2022 · 194 citations
- Capsule-based Object Tracking with Natural Language SpecificationDing Ma, Xiangqian WuACM MM 2021 · 25 citations
- Employing Deep Part-Object Relationships for Salient Object DetectionYi Liu, Qiang Zhang, Dingwen Zhang, Jungong HanICCV 2019 · 86 citations
