A Closer Look at Self-Supervised Lightweight Vision Transformers
Shaoru Wang, Jin Gao, Zeming Li, Xiaoqin Zhang, Weiming Hu
Abstract
Self-supervised learning on large-scale Vision Transformers (ViTs) as pre-training methods has achieved promising downstream performance. Yet, how much these pre-training paradigms promote lightweight ViTs' performance is considerably less studied. In this work, we develop and benchmark several self-supervised pre-training methods on image classification tasks and some downstream dense prediction tasks. We surprisingly find that if proper pre-training is adopted, even vanilla lightweight ViTs show comparable performance to previous SOTA networks with delicate architecture design. It breaks the recently popular conception that vanilla ViTs are not suitable for vision tasks in lightweight regimes. We also point out some defects of such pre-training, e.g., failing to benefit from large-scale pre-training data and showing inferior performance on data-insufficient downstream tasks. Furthermore, we analyze and clearly show the effect of such pre-training by analyzing the properties of the layer representation and attention maps for related models. Finally, based on the above analyses, a distillation strategy during pre-training is developed, which leads to further downstream performance improvement for MAE-based pre-training. Code is available at https://github.com/wangsr126/mae-lite .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0c73fe9-9759-4014-a806-a51636b0e30aCited by top-tier papers16
- ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual TrackingYutong Kou, Jin Gao, Bing Li, Gang Wang et al.NeurIPS 2023 · 74 citations
- What Do Self-Supervised Vision Transformers Learn?Namuk Park, Wonjae Kim, Byeongho Heo, Taekyung Kim et al.ICLR 2023 · 16 citations
- Masked Image Residual Learning for Scaling Deeper Vision TransformersGuoxi Huang, Hongtao Fu, Adrian G. BorsNeurIPS 2023 · 10 citations
- Hybrid Distillation: Connecting Masked Autoencoders with Contrastive LearnersBowen Shi, Xiaopeng Zhang, Yaoming Wang, Jin Li et al.ICLR 2024 · 10 citations
- MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data UtilizationYu Zhang, Qi Zhang, Zixuan Gong, Yiwei Shi et al.ICML 2024 · 9 citations
Builds on34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Closest Neighbors are Harmful for Lightweight Masked Auto-encodersJian Meng, Ahmed Hassan, Li Yang, Deliang Fan et al.CVPR 2025
- Generic-to-Specific Distillation of Masked AutoencodersWei Huang, Zhiliang Peng, Li Dong, Furu Wei et al.CVPR 2023
- On the Surprising Effectiveness of Attention Transfer for Vision TransformersAlexander C. Li, Yuandong Tian, Beidi Chen, Deepak Pathak et al.NeurIPS 2024 · 21 citations
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 52 citations
- SparseMAE: Sparse Training Meets Masked AutoencodersAojun Zhou, Yang Li, Zipeng Qin, Jianbo Liu et al.ICCV 2023 · 8 citations
