Lune

CVPR2020Top-tier venue

Deformable Siamese Attention Networks for Visual Object Tracking

Yuechen Yu, Yilei Xiong, Weilin Huang, Matthew R. Scott

2020Year
38Top-tier citations

Abstract

Siamese-based trackers have achieved excellent performance on visual object tracking. However, the target template is not updated online, and the features of the target template and search image are computed independently in a Siamese architecture. In this paper, we propose Deformable Siamese Attention Networks, referred to as Sia-mAttn, by introducing a new Siamese attention mechanism that computes deformable self-attention and crossattention. The self-attention learns strong context information via spatial attention, and selectively emphasizes interdependent channel-wise features with channel attention. The cross-attention is capable of aggregating rich contextual interdependencies between the target template and the search image, providing an implicit manner to adaptively update the target template. In addition, we design a region refinement module that computes depth-wise cross correlations between the attentional features for more accurate tracking. We conduct experiments on six benchmarks, where our method achieves new state-of-the-art results, outperforming the strong baseline, SiamRPN++ [24], by 0.464→0.537 and 0.415→0.470 EAO on VOT 2016 and 2018. Our code is available at: https://github.com/msight- tech/research-siamattn.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 19313b3d-43ee-44bf-8220-03cb28c81eda

Cited by top-tier papers38

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines