Self-supervised Models are Good Teaching Assistants for Vision Transformers
Haiyan Wu, Yuting Gao, Yinqi Zhang, Shaohui Lin, Yuan Xie, Xing Sun, Ke Li
Abstract
Transformers have shown remarkable progress on computer vision tasks in the past year. Compared to their CNN counterparts, transformers usually need the help of distillation to achieve comparable results on middle or small sized datasets. Meanwhile, recent researches discover that when transformers are trained with supervised and self-supervised manner respectively, the captured patterns are quite different both qualitatively and quantitatively. These findings motivate us to introduce a self-supervised teaching assistant (SSTA) besides the commonly used supervised teacher to improve the performance of transformers. Specifically, we propose a headlevel knowledge distillation method that selects the most important head of the supervised teacher and self-supervised teaching assistant, and let the student mimic the attention distribution of these two heads, so as to make the student focus on the relationship between tokens deemed by the teacher and the teacher assistant. Extensive experiments verify the effectiveness of SSTA and demonstrate that the proposed SSTA is a good compensation to the supervised teacher. Meanwhile, some analytical experiments towards multiple perspectives (e.g. prediction, shape bias, robustness, and transferability to downstream tasks) with supervised teachers, self-supervised teaching assistants and students are inductive and may inspire future researches. The code is released in https://github.com/GlassyWu/SSTA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed612160-daf9-4fba-992b-7a8cef2a01a4Cited by top-tier papers5
- EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment AnythingYunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xiang et al.CVPR 2024 · 185 citations
- CDAC: Cross-domain Attention Consistency in Transformer for Domain Adaptive Semantic SegmentationKaihong Wang, Donghyun Kim, Rogério Feris, Margrit BetkeICCV 2023 · 32 citations
- Asymmetric Masked Distillation for Pre-Training Small Foundation ModelsZhiyu Zhao, Bingkun Huang, Sen Xing, Gangshan Wu et al.CVPR 2024 · 6 citations
- Boosting Vanilla Lightweight Vision Transformers via Re-parameterizationZhentao Tan, Xiaodan Li, Yue Wu, Qi Chu et al.ICLR 2024 · 5 citations
- Generic-to-Specific Distillation of Masked AutoencodersWei Huang, Zhiliang Peng, Li Dong, Furu Wei et al.CVPR 2023
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Co-advise: Cross Inductive Bias DistillationSucheng Ren, Zhengqi Gao, Tianyu Hua, Zihui Xue et al.CVPR 2022 · 50 citations
- Knowledge Distillation via the Target-aware TransformerSihao Lin, Hongwei Xie, Bing Wang, Kaicheng Yu et al.CVPR 2022 · 126 citations
- Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen et al.CVPR 2023
- A Closer Look at Self-Supervised Lightweight Vision TransformersShaoru Wang, Jin Gao, Zeming Li, Xiaoqin Zhang et al.ICML 2023 · 61 citations
- Supervised Masked Knowledge Distillation for Few-Shot TransformersHan Lin, Guangxing Han, Jiawei Ma, Shiyuan Huang et al.CVPR 2023
