Feature Kernel Distillation
Bobby He, Mete Ozay
摘要
Trained Neural Networks (NNs) can be viewed as data-dependent kernel machines, with predictions determined by the inner product of last-layer representations across inputs, referred to as the feature kernel. We explore the relevance of the feature kernel for Knowledge Distillation (KD), using a mechanistic understanding of an NN's optimisation process. We extend the theoretical analysis of Allen-Zhu & Li (2020) to show that a trained NN's feature kernel is highly dependent on its parameter initialisation, which biases different initialisations of the same architecture to learn different data attributes in a multi-view data setting. This enables us to prove that KD using only pairwise feature kernel comparisons can improve NN test accuracy in such settings, with both single & ensemble teacher models, whereas standard training without KD fails to generalise. We further use our theory to motivate practical considerations for improving student generalisation when using distillation with feature kernels, which allows us to propose a novel approach: Feature Kernel Distillation (FKD). Finally, we experimentally corroborate our theory in the image classification setting, showing that FKD is amenable to ensemble distillation, can transfer knowledge across datasets, and outperforms both vanilla KD & other feature kernel based KD baselines across a range of standard architectures & datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationYichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang 等NeurIPS 2024 · 被引用 94 次
- Teach Less, Learn More: On the Undistillable Classes in Knowledge DistillationYichen Zhu, Ning Liu, Zhiyuan Xu, Xin Liu 等NeurIPS 2022 · 被引用 42 次
- Bring Evanescent Representations to Life in Lifelong Class Incremental LearningMarco Toldo, Mete OzayCVPR 2022 · 被引用 34 次
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag 等NeurIPS 2024 · 被引用 27 次
- : Improving Knowledge Distillation Using Orthogonal ProjectionsRoy Miles, Ismail Elezi, Jiankang DengCVPR 2024 · 被引用 9 次
它引用的顶会 Paper18
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Cross-Layer Distillation with Semantic CalibrationDefang Chen, Jian-Ping Mei, Yuan Zhang, Can Wang 等AAAI 2021 · 被引用 368 次
相关 Paper
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- Improved Feature Distillation via Projector EnsembleYudong Chen, Sen Wang, Jiajun Liu, Xuwei Xu 等NeurIPS 2022 · 被引用 73 次
- Distilling Knowledge from Heterogeneous Architectures for Semantic SegmentationYanglin Huang, Kai Hu, Yuan Zhang, Zhineng Chen 等AAAI 2025 · 被引用 4 次
- Revisiting Knowledge Distillation: An Inheritance and Exploration FrameworkZhen Huang, Xu Shen, Jun Xing, Tongliang Liu 等CVPR 2021
- Search to Distill: Pearls Are Everywhere but Not the EyesYu Liu, Xuhui Jia, Mingxing Tan, Raviteja Vemulapalli 等CVPR 2020
