MAST: A Memory-Augmented Self-Supervised Tracker
Zihang Lai, Erika Lu, Weidi Xie
Abstract
Recent interest in self-supervised dense tracking has yielded rapid progress, but performance still remains far from supervised methods. We propose a dense tracking model trained on videos without any annotations that surpasses previous self-supervised methods on existing benchmarks by a significant margin (+15%), and achieves performance comparable to supervised methods. In this paper, we first reassess the traditional choices used for selfsupervised training and reconstruction loss by conducting thorough experiments that finally elucidate the optimal choices. Second, we further improve on existing methods by augmenting our architecture with a crucial memory component. Third, we benchmark on large-scale semi-supervised video object segmentation (aka. dense tracking), and propose a new metric: generalizability. Our first two contributions yield a self-supervised network that for the first time is competitive with supervised methods on standard evaluation metrics of dense tracking. When measuring generalizability, we show self-supervised approaches are actually superior to the majority of supervised methods. We believe this new generalizability metric can better capture the realworld use-cases for dense tracking, and will spur new interest in this research direction. Code will be released at https://github.com/zlai0/MAST .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a4a6a99-5848-4d97-9e9f-67e8971c29c9Cited by top-tier papers63
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Emergent Correspondence from Image DiffusionLuming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo et al.NeurIPS 2023 · 555 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra et al.NeurIPS 2021 · 382 citations
Builds on2
Related papers
- Video Diffusion Models Excel at Tracking Similar-Looking Objects Without SupervisionChenshuang Zhang, Kang Zhang, Joon Son Chung, In So Kweon et al.NeurIPS 2025
- TrackMAE: Video Representation Learning via Track Mask and PredictRenaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard GhanemCVPR 2026 · 3 citations
- Self-Supervised Multi-Object Tracking with Cross-input ConsistencyFavyen Bastani, Songtao He, Samuel MaddenNeurIPS 2021 · 39 citations
- Learning to Track Instances without Video AnnotationsYang Fu, Sifei Liu, Umar Iqbal, Shalini De Mello et al.CVPR 2021
- Spatiotemporal Graph Neural Network based Mask Reconstruction for Video Object SegmentationDaizong Liu, Shuangjie Xu, Xiao-Yang Liu, Zichuan Xu et al.AAAI 2021 · 25 citations
