Positional Label for Self-Supervised Vision Transformer
Zhemin Zhang, Xun Gong
Abstract
Positional encoding is important for vision transformer (ViT) to capture the spatial structure of the input image. General effectiveness has been proven in ViT. In our work we propose to train ViT to recognize the positional label of patches of the input image, this apparently simple task actually yields a meaningful self-supervisory task. Based on previous work on ViT positional encoding, we propose two positional labels dedicated to 2D images including absolute position and relative position. Our positional labels can be easily plugged into various current ViT variants. It can work in two ways: (a) As an auxiliary training target for vanilla ViT (e.g., ViT-B [Dosovitskiy et al., 2020] and Swin-B [Liu et al., 2021]) for better performance. (b) Combine the self-supervised ViT (e.g., MAE [He et al., 2021]) to provide a more powerful self-supervised signal for semantic feature learning. Experiments demonstrate that with the proposed self-supervised methods, ViT-B and Swin-B gain improvements of 1.20% (top-1 Acc) and 0.74% (top-1 Acc) on ImageNet, respectively, and 6.15% and 1.14% improvement on Mini-ImageNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ff2dbb8-2b4c-4d17-830f-b862bf64690eCited by top-tier papers2
- VoCo: A Simple-Yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisLinshan Wu, Jiaxin Zhuang, Hao ChenCVPR 2024 · 60 citations
- Structure-Aware Semantic Discrepancy and Consistency for 3D Medical Image Self-Supervised LearningTan Pan, Zhaorui Tan, Kaiyu Guo, Dongli Xu et al.ICCV 2025 · 2 citations
Builds on7
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- HiViT: A Simpler and More Efficient Design of Hierarchical Vision TransformerXiaosong Zhang, Yunjie Tian, Lingxi Xie, Wei Huang et al.ICLR 2023
- All Tokens Matter: Token Labeling for Training Better Vision TransformersZihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou et al.NeurIPS 2021 · 252 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- DropPos: Pre-Training Vision Transformers by Reconstructing Dropped PositionsHaochen Wang, Junsong Fan, Yuxi Wang, Kaiyou Song et al.NeurIPS 2023 · 32 citations
- LaPE: Layer-adaptive Position Embedding for Vision Transformers with Independent Layer NormalizationRunyi Yu, Zhennan Wang, Yinhuai Wang, Kehan Li et al.ICCV 2023 · 13 citations
