Co-advise: Cross Inductive Bias Distillation
Sucheng Ren, Zhengqi Gao, Tianyu Hua, Zihui Xue, Yonglong Tian, Shengfeng He, Hang Zhao
Abstract
The inductive bias of vision transformers is more relaxed that cannot work well with insufficient data. Knowledge distillation is thus introduced to assist the training of transformers. Unlike previous works, where merely heavy convolution-based teachers are provided, in this paper, we delve into the influence of models inductive biases in knowledge distillation (e.g., convolution and involution). Our key observation is that the teacher accuracy is not the dominant reason for the student accuracy, but the teacher inductive bias is more important. We demonstrate that lightweight teachers with different architectural inductive biases can be used to co-advise the student transformer with outstanding performances. The rationale behind is that models designed with different inductive biases tend to focus on diverse patterns, and teachers with different inductive biases attain various knowledge despite being trained on the same dataset. The diverse knowledge provides a more precise and comprehensive description of the data and compounds and boosts the performance of the student during distillation. Furthermore, we propose a token inductive bias alignment to align the inductive bias of the token with its target teacher model. With only lightweight teachers provided and using this cross inductive bias distillation method, our vision transformers (termed as CiT) outperform all previous vision transformers (ViT) of the same architecture on ImageNet. Moreover, our small size model CiT-SAK further achieves 82.7% Top-1 accuracy on ImageNet without modifying the attention module of the ViT. Code is available at https://github.com/OliverRensu/co-advise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76a4b00a-03dd-4917-84cd-b8a0408319fcCited by top-tier papers13
- SG-Former: Self-guided Transformer with Evolving Token ReallocationSucheng Ren, Xingyi Yang, Songhua Liu, Xinchao WangICCV 2023 · 70 citations
- Cumulative Spatial Knowledge Distillation for Vision TransformersBorui Zhao, Renjie Song, Jiajun LiangICCV 2023 · 27 citations
- Revisit the Power of Vanilla Knowledge Distillation: from Small Scale to Large ScaleZhiwei Hao, Jianyuan Guo, Kai Han, Han Hu et al.NeurIPS 2023 · 17 citations
- The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge DistillationZihui Xue, Zhengqi Gao, Sucheng Ren, Hang ZhaoICLR 2023 · 12 citations
- D3still: Decoupled Differential Distillation for Asymmetric Image RetrievalYi Xie, Yihong Lin, Wenjie Cai, Xuemiao Xu et al.CVPR 2024 · 10 citations
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
- Generic-to-Specific Distillation of Masked AutoencodersWei Huang, Zhiliang Peng, Li Dong, Furu Wei et al.CVPR 2023
- DeiT-LT: Distillation Strikes Back for Vision Transformer Training on Long-Tailed DatasetsHarsh Rangwani, Pradipto Mondal, Mayank Mishra, Ashish Ramayee Asokan et al.CVPR 2024 · 12 citations
- Bootstrapping ViTs: Towards Liberating Vision Transformers from Pre-trainingHaofei Zhang, Jiarui Duan, Mengqi Xue, Jie Song et al.CVPR 2022 · 19 citations
- TinyMIM: An Empirical Study of Distilling MIM Pre-trained ModelsSucheng Ren, Fangyun Wei, Zheng Zhang, Han HuCVPR 2023
