: Improving Knowledge Distillation Using Orthogonal Projections
Roy Miles, Ismail Elezi, Jiankang Deng
Abstract
Knowledge distillation is an effective method for training small and efficient deep learning models. However, the efficacy of a single method can degenerate when transferring to other tasks, modalities, or even other architectures. To address this limitation, we propose a novel constrained feature distillation method. This method is derived from a small set of core principles, which results in two emerging components: an orthogonal projection and a task-specific normalisation. Equipped with both of these components, our transformer models can outperform all previous methods on ImageNet and reach up to a 4.4% relative improvement over the previous state-of-the-art methods. To further demonstrate the generality of our method, we apply it to object detection and image generation, whereby we obtain consistent and substantial performance improvements over state-of-the-art. Code and models are publicly available<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>https://github.com/roymiles/vkd.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8e1da8d-3b2c-4a8f-ba35-1dc6f5ac0a23Cited by top-tier papers10
- Feature Distillation is the Better Choice for Model-Heterogeneous Federated LearningYichen Li, Xiuying Wang, Wenchao Xu, Haozhao Wang et al.NeurIPS 2025 · 6 citations
- Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary HeadPenghui Yang, Chen-Chen Zong, Sheng-Jun Huang, Lei Feng et al.KDD 2025 · 1 citation
- SRA: Span Representation Alignment for Large Language Model DistillationQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Tung Nguyen et al.ACL 2026 · 1 citation
- All You Need in Knowledge Distillation Is a Tailored Coordinate SystemJunjie Zhou, Ke Zhu, Jianxin WuAAAI 2025 · 1 citation
- Late-to-Early Training: LET LLMs Learn Earlier, So Faster and BetterJi Zhao, Shitong Shao, Yufei Gu, Xun Zhou et al.ICLR 2026 · 1 citation
Builds on36
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
Related papers
- ScaleKD: Distilling Scale-Aware Knowledge in Small Object DetectorYichen Zhu, Qiqi Zhou, Ning Liu, Zhiyuan Xu et al.CVPR 2023
- Generic-to-Specific Distillation of Masked AutoencodersWei Huang, Zhiliang Peng, Li Dong, Furu Wei et al.CVPR 2023
- Task-Oriented Feature DistillationLinfeng Zhang, Yukang Shi, Zuoqiang Shi, Kaisheng Ma et al.NeurIPS 2020 · 74 citations
- Distilling DETR with Visual-Linguistic Knowledge for Open-Vocabulary Object DetectionLiangqi Li, Jiaxu Miao, Dahu Shi, Wenming Tan et al.ICCV 2023 · 35 citations
- Are Large Kernels Better Teachers than Transformers for ConvNets?Tianjin Huang, Lu Yin, Zhenyu Zhang, Li Shen et al.ICML 2023 · 18 citations
