Teaching Matters: Investigating the Role of Supervision in Vision Transformers
Matthew Walmer, Saksham Suri, Kamal Gupta, Abhinav Shrivastava
Abstract
Vision Transformers (ViTs) have gained significant popularity in recent years and have proliferated into many applications. However, their behavior under different learning paradigms is not well explored. We compare ViTs trained through different methods of supervision, and show that they learn a diverse range of behaviors in terms of their attention, representations, and downstream performance. We also discover ViT behaviors that are consistent across supervision, including the emergence of Offset Local Attention Heads. These are self-attention heads that attend to a token adjacent to the current token with a fixed directional offset, a phenomenon that to the best of our knowledge has not been highlighted in any prior work. Our analysis shows that ViTs are highly flexible and learn to process local and global information in different orders depending on their training method. We find that contrastive self-supervised methods learn features that are competitive with explicitly supervised features, and they can even be superior for part-level tasks. We also find that the representations of reconstruction-based models show non-trivial similarity to contrastive self-supervised models. Project website and code are publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9accdc57-4397-42cd-a5be-c63e4378ea90Cited by top-tier papers20
- Perception Encoder: The best visual embeddings are not at the output of the networkDaniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho et al.NeurIPS 2025 · 359 citations
- MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoChenjie Cao, Xinlin Ren, Yanwei FuICLR 2024 · 68 citations
- Contrastive Tuning: A Little Help to Make Masked Autoencoders ForgetJohannes Lehner, Benedikt Alkin, Andreas Fürst, Elisabeth Rumetshofer et al.AAAI 2024 · 28 citations
- ConvNet vs Transformer, Supervised vs CLIP: Beyond ImageNet AccuracyKirill Vishniakov, Zhiqiang Shen, Zhuang LiuICML 2024 · 26 citations
- On the Surprising Effectiveness of Attention Transfer for Vision TransformersAlexander C. Li, Yuandong Tian, Beidi Chen, Deepak Pathak et al.NeurIPS 2024 · 21 citations
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- A Theoretical Analysis of Self-Supervised Learning for Vision TransformersYu Huang, Zixin Wen, Yuejie Chi, Yingbin LiangICLR 2025
- Vision Transformers Need More Than RegistersCheng Shi, Yizhou Yu, Sibei YangCVPR 2026 · 17 citations
- What Do Self-Supervised Vision Transformers Learn?Namuk Park, Wonjae Kim, Byeongho Heo, Taekyung Kim et al.ICLR 2023 · 16 citations
- Vision Transformers provably learn spatial structureSamy Jelassi, Michael E. Sander, Yuanzhi LiNeurIPS 2022 · 115 citations
- Self-Supervised Learning of Intertwined Content and Positional Features for Object DetectionKang-Jun Liu, Masanori Suganuma, Takayuki OkataniICML 2025
