Search to Distill: Pearls Are Everywhere but Not the Eyes
Yu Liu, Xuhui Jia, Mingxing Tan, Raviteja Vemulapalli, Yukun Zhu, Bradley Green, Xiaogang Wang
摘要
Standard Knowledge Distillation (KD) approaches distill the knowledge of a cumbersome teacher model into the parameters of a student model with a pre-defined architecture. However, the knowledge of a neural network, which is represented by the network's output distribution conditioned on its input, depends not only on its parameters but also on its architecture. Hence, a more generalized approach for KD is to distill the teacher's knowledge into both the parameters and architecture of the student. To achieve this, we present a new Architecture-aware Knowledge Distillation (AKD) approach that finds student models (pearls for the teacher) that are best for distilling the given teacher model. In particular, we leverage Neural Architecture Search (NAS), equipped with our KD-guided reward, to search for the best student architectures for a given teacher. Experimental results show our proposed AKD consistently outperforms the conventional NAS plus KD approach, and achieves state-of-the-art results on the ImageNet classification task under various latency settings. Furthermore, the best AKD student architecture for the ImageNet classification task also transfers well to other tasks such as million level face recognition and ensemble learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 被引用 103 次
- Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture SearchHouwen Peng, Hao Du, Hongyuan Yu, Qi Li 等NeurIPS 2020 · 被引用 76 次
- Task-Oriented Feature DistillationLinfeng Zhang, Yukang Shi, Zuoqiang Shi, Kaisheng Ma 等NeurIPS 2020 · 被引用 74 次
- KD-Zero: Evolving Knowledge Distiller for Any Teacher-Student PairsLujun Li, Peijie Dong, Anggeng Li, Zimian Wei 等NeurIPS 2023 · 被引用 49 次
- NORM: Knowledge Distillation via N-to-One Representation MatchingXiaolong Liu, Lujun Li, Chao Li, Anbang YaoICLR 2023 · 被引用 19 次
它引用的顶会 Paper3
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou 等ICCV 2019 · 被引用 625 次
- Knowledge Distillation via Route Constrained OptimizationXiao Jin, Baoyun Peng, Yichao Wu, Yu Liu 等ICCV 2019 · 被引用 196 次
- Differentiable Kernel EvolutionYu Liu, Jihao Liu, Xiaogang Wang, Ailing ZengICCV 2019 · 被引用 6 次
相关 Paper
- Teacher Guided Neural Architecture Search for Face RecognitionXiaobo WangAAAI 2021 · 被引用 12 次
- Towards Oracle Knowledge Distillation with Neural Architecture SearchMinsoo Kang, Jonghwan Mun, Bohyung HanAAAI 2020 · 被引用 48 次
- Meta-prediction Model for Distillation-Aware NAS on Unseen DatasetsHayeon Lee, Sohyun An, Minseon Kim, Sung Ju HwangICLR 2023 · 被引用 2 次
- Performance-Aware Mutual Knowledge Distillation for Improving Neural Architecture SearchPengtao Xie, Xuefeng DuCVPR 2022 · 被引用 16 次
- Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language ModelsDongkuan Xu, Subhabrata Mukherjee, Xiaodong Liu, Debadeepta Dey 等NeurIPS 2022 · 被引用 21 次
