VideoTrack: Learning to Track Objects via Video Transformer
Fei Xie, Lei Chu, Jiahao Li, Yan Lu, Chao Ma
摘要
Existing Siamese tracking methods, which are built on pair-wise matching between two single frames, heavily rely on additional sophisticated mechanism to exploit temporal information among successive video frames, hindering them from efficiency and industrial deployments. In this work, we resort to sequence-level target matching that can encode temporal contexts into the spatial features through a neat feedforward video model. Specifically, we adapt the standard video transformer architecture to visual tracking by enabling spatiotemporal feature learning directly from frame-level patch sequences. To better adapt to the tracking task, we carefully blend the spatiotemporal information in the video clips through sequential multi-branch triplet blocks, which formulates a video transformer backbone. Our experimental study compares different model variants, such as tokenization strategies, hierarchical structures, and video attention schemes. Then, we propose a disentangled dual-template mechanism that decouples static and dynamic appearance clues over time, and reduces temporal redundancy in video frames. Extensive experiments show that our method, named as VideoTrack, achieves state-ofthe-art results while running in real-time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo 等AAAI 2024 · 被引用 247 次
- Explicit Visual Prompts for Visual Object TrackingLiangtao Shi, Bineng Zhong, Qihua Liang, Ning Li 等AAAI 2024 · 被引用 117 次
- Autoregressive Queries for Adaptive Tracking with Spatio-Temporal TransformersJinxia Xie, Bineng Zhong, Zhiyi Mo, Shengping Zhang 等CVPR 2024 · 被引用 100 次
- Exploring Enhanced Contextual Information for Video-Level Object TrackingBen Kang, Xin Chen, Simiao Lai, Yang Liu 等AAAI 2025 · 被引用 48 次
- Decoupled Spatio-Temporal Consistency Learning for Self-Supervised TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Ning Li 等AAAI 2025 · 被引用 41 次
它引用的顶会 Paper33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual TrackingNing Wang, Wengang Zhou, Jie Wang, Houqiang LiCVPR 2021
- SeqTrack: Sequence to Sequence Learning for Visual Object TrackingXin Chen, Houwen Peng, Dong Wang, Huchuan Lu 等CVPR 2023
- An Efficient Token Compression Framework for Visual Object TrackingWeijing Wu, Qihua Liang, Bineng Zhong, Haiying Xia 等CVPR 2026 · 被引用 1 次
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 被引用 16 次
- Fast Spatial Tracking with Visual Geometry TransformerChengjie Huang, GUILE WU, Dongfeng Bai, Bingbing LiuCVPR 2026
