Lune

CVPR2023Top-tier venue

SeqTrack: Sequence to Sequence Learning for Visual Object Tracking

Xin Chen, Houwen Peng, Dong Wang, Huchuan Lu, Han Hu

2023Year
82Top-tier citations

Abstract

In this paper, we present a new sequence-to-sequence learning framework for visual tracking, dubbed SeqTrack. It casts visual tracking as a sequence generation problem, which predicts object bounding boxes in an autoregressive fashion. This is different from prior Siamese trackers and transformer trackers, which rely on designing complicated head networks, such as classification and regression heads. SeqTrack only adopts a simple encoder-decoder transformer architecture. The encoder extracts visual features with a bidirectional transformer, while the decoder generates a sequence of bounding box values autoregressively with a causal transformer. The loss function is a plain cross-entropy. Such a sequence learning paradigm not only simplifies tracking framework, but also achieves competitive performance on benchmarks. For instance, Se-qTrack gets 72.5% AUC on LaSOT, establishing a new stateof-the-art performance. Code and models are available at https://github.com/microsoft/VideoX .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c8b646c5-c0e1-450d-8369-1628cdda03e0

Cited by top-tier papers82

Ask how each one uses it

Builds on25

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines