One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation
Zhiwei Hao, Jianyuan Guo, Kai Han, Yehui Tang, Han Hu, Yunhe Wang, Chang Xu
Abstract
Knowledge distillation (KD) has proven to be a highly effective approach for enhancing model performance through a teacher-student training scheme. However, most existing distillation methods are designed under the assumption that the teacher and student models belong to the same model family, particularly the hint-based approaches. By using centered kernel alignment (CKA) to compare the learned features between heterogeneous teacher and student models, we observe significant feature divergence. This divergence illustrates the ineffectiveness of previous hint-based methods in cross-architecture distillation. To tackle the challenge in distilling heterogeneous models, we propose a simple yet effective one-for-all KD framework called OFA-KD, which significantly improves the distillation performance between heterogeneous architectures. Specifically, we project intermediate features into an aligned latent space such as the logits space, where architecture-specific information is discarded. Additionally, we introduce an adaptive target enhancement scheme to prevent the student from being disturbed by irrelevant information. Extensive experiments with various architectures, including CNN, Transformer, and MLP, demonstrate the superiority of our OFA-KD framework in enabling distillation between heterogeneous architectures. Specifically, when equipped with our OFA-KD, the student models achieve notable performance improvements, with a maximum gain of 8.0% on the CIFAR-100 dataset and 0.7% on the ImageNet-1K dataset. PyTorch code 1 and checkpoints can be found at https://github.com/Hao840/OFAKD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2fbac82-947c-4564-a54b-675d27e73e1cCited by top-tier papers36
- Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge DistillationJiaming Lv, Haoyuan Yang, Peihua LiNeurIPS 2024 · 59 citations
- ScaleKD: Strong Vision Transformers Could Be Excellent TeachersJiawei Fan, Chao Li, Xiaolong Liu, Anbang YaoNeurIPS 2024 · 20 citations
- DistillDrive: End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning ModelRui Yu, Xianghang Zhang, Runkai Zhao, Huaicheng Yan et al.ICCV 2025 · 19 citations
- Task Groupings Regularization: Data-Free Meta-Learning with Heterogeneous Pre-trained ModelsYongxian Wei, Zixuan Hu, Li Shen, Zhenyi Wang et al.ICML 2024 · 11 citations
- Local Dense Logit Relations for Enhanced Knowledge DistillationLiuchi Xu, Kang Liu, Jinshuai Liu, Lu Wang et al.ICCV 2025 · 10 citations
Builds on33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Distilling Knowledge from Heterogeneous Architectures for Semantic SegmentationYanglin Huang, Kai Hu, Yuan Zhang, Zhineng Chen et al.AAAI 2025 · 4 citations
- Fuse Before Transfer: Knowledge Fusion for Heterogeneous DistillationGuopeng Li, Qiang Wang, Ke Yan, Shouhong Ding et al.ICCV 2025 · 1 citation
- Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous DistillationJhe-Hao Lin, Yi Yao, Chan-Feng Hsu, Hong-Xia Xie et al.ICCV 2025 · 3 citations
- Cross-Architecture Distillation Made Simple with Redundancy SuppressionWeijia Zhang, Yuehao Liu, Wu Ran, Chao MaICCV 2025 · 6 citations
- Heterogeneous Complementary DistillationLiuchi Xu, Hao Zheng, Lu Wang, Lisheng Xu et al.AAAI 2026
