Autoregressive Sequential Pretraining for Visual Tracking
Shiyi Liang, Yifan Bai, Yihong Gong, Xing Wei
摘要
Recent advancements in visual object tracking have shifted towards a sequential generation paradigm, where object deformation and motion exhibit strong temporal dependencies. Despite the importance of these dependencies, widely adopted image-level pretrained backbones barely capture the dynamics in the consecutive video, which is the essence of tracking. Thus, we propose AutoRegressive Sequential Pretraining (ARP), an unsupervised spatio-temporal learner, via generating the evolution of object appearance and motion in video sequences. Our method leverages a diffusion model to autoregressively generate the future frame appearance, conditioned on historical embeddings extracted by a general encoder. Furthermore, to ensure trajectory coherence, the same encoder is employed to learn trajectory consistency by generating coordinate sequences in a reverse autoregressive fashion, a process we term backtracking. Further, we integrate the pretrained ARP into AR-TrackV2, creating ARPTrack, which is further fine-tuned for tracking tasks. ARPTrack achieves state-of-the-art performance across multiple benchmarks, becoming the first tracker to surpass 80% AO on GOT-10k, while maintaining high efficiency. These results demonstrate the effectiveness of our approach in capturing temporal dependencies for continuous video tracking. The code will be released soon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- FARTrack: Fast Autoregressive Visual Tracking with High PerformanceGuijie Wang, Tong Lin, Yifan Bai, Anjia Cao 等ICLR 2026 · 被引用 3 次
- An Efficient Token Compression Framework for Visual Object TrackingWeijing Wu, Qihua Liang, Bineng Zhong, Haiying Xia 等CVPR 2026 · 被引用 1 次
- RELO: Reinforcement Learning to Localize for Visual Object TrackingXin Chen, Chuanyu Sun, Jiao Xu, Houwen Peng 等ICML 2026 · 被引用 1 次
- GOT-Edit: Geometry-Aware Generic Object Tracking via Online Model EditingShih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu LinICLR 2026 · 被引用 1 次
- Boosting Self-Supervised Tracking with Contextual Prompts and Noise LearningYaozong Zheng, Qihua Liang, Bineng Zhong, Shuimu Zeng 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper37
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
相关 Paper
- ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to DescribeYifan Bai, Zeyang Zhao, Yihong Gong, Xing WeiCVPR 2024
- Autoregressive Visual TrackingXing Wei, Yifan Bai, Yongchao Zheng, Dahu Shi 等CVPR 2023
- TGTrack: Temporal Generative Learning for Unified Single Object TrackingWanting Geng, Xin Chen, Chuanyu Sun, Jie Zhao 等CVPR 2026
- SeqTrack: Sequence to Sequence Learning for Visual Object TrackingXin Chen, Houwen Peng, Dong Wang, Huchuan Lu 等CVPR 2023
- ODTrack: Online Dense Temporal Token Learning for Visual TrackingYaozong Zheng, Bineng Zhong, Qihua Liang, Zhiyi Mo 等AAAI 2024 · 被引用 247 次
