Representation Learning for Visual Object Tracking by Masked Appearance Transfer
Haojie Zhao, Dong Wang, Huchuan Lu
Abstract
Visual representation plays an important role in visual object tracking. However, few works study the trackingspecified representation learning method. Most trackers directly use ImageNet pre-trained representations. In this paper, we propose masked appearance transfer, a simple but effective representation learning method for tracking, based on an encoder-decoder architecture. First, we encode the visual appearances of the template and search region jointly, and then we decode them separately. During decoding, the original search region image is reconstructed. However, for the template, we make the decoder reconstruct the target appearance within the search region. By this target appearance transfer, the tracking-specified representations are learned. We randomly mask out the inputs, thereby making the learned representations more discriminative. For sufficient evaluation, we design a simple and lightweight tracker that can evaluate the representation for both target localization and box regression. Extensive experiments show that the proposed method is effective, and the learned representations can enable the simple tracker to obtain state-of-the-art performance on six datasets. https://github.com/difhnp/MAT
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26c507dc-1426-4c55-92e6-bb3383e1a7c8Cited by top-tier papers12
- Single-Model and Any-Modality for Video Object TrackingZongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu et al.CVPR 2024 · 78 citations
- Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV TrackingYongxin Li, Mengyuan Liu, You Wu, Xucheng Wang et al.ICML 2024 · 63 citations
- ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelYiming Sun, Fan Yu, Shaoxiang Chen, Yu Zhang et al.NeurIPS 2024 · 21 citations
- DeTrack: In-model Latent Denoising Learning for Visual Object TrackingXinyu Zhou, Jinglun Li, Lingyi Hong, Kaixun Jiang et al.NeurIPS 2024 · 14 citations
- Alligat0R: Pre-Training through Covisibility Segmentation for Relative Camera Pose RegressionThibaut Loiseau, Guillaume Bourmaud, Vincent LepetitNeurIPS 2025 · 11 citations
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
Related papers
- Exploring Target Representations for Masked AutoencodersXingbin Liu, Jinghao Zhou, Tao Kong, Xianming Lin et al.ICLR 2024 · 59 citations
- Optical Flow in Deep Visual TrackingMikko Vihlman, Arto VisalaAAAI 2020 · 17 citations
- What Makes Instance Discrimination Good for Transfer Learning?Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen LinICLR 2021 · 183 citations
- Aligning Pretraining for Detection via Object-Level Contrastive LearningFangyun Wei, Yue Gao, Zhirong Wu, Han Hu et al.NeurIPS 2021 · 180 citations
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
