Search to Distill: Pearls Are Everywhere but Not the Eyes
Yu Liu, Xuhui Jia, Mingxing Tan, Raviteja Vemulapalli, Yukun Zhu, Bradley Green, Xiaogang Wang
Abstract
Standard Knowledge Distillation (KD) approaches distill the knowledge of a cumbersome teacher model into the parameters of a student model with a pre-defined architecture. However, the knowledge of a neural network, which is represented by the network's output distribution conditioned on its input, depends not only on its parameters but also on its architecture. Hence, a more generalized approach for KD is to distill the teacher's knowledge into both the parameters and architecture of the student. To achieve this, we present a new Architecture-aware Knowledge Distillation (AKD) approach that finds student models (pearls for the teacher) that are best for distilling the given teacher model. In particular, we leverage Neural Architecture Search (NAS), equipped with our KD-guided reward, to search for the best student architectures for a given teacher. Experimental results show our proposed AKD consistently outperforms the conventional NAS plus KD approach, and achieves state-of-the-art results on the ImageNet classification task under various latency settings. Furthermore, the best AKD student architecture for the ImageNet classification task also transfers well to other tasks such as million level face recognition and ensemble learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eef3a5ba-73bf-4bd5-8707-9ba7e020c833Cited by top-tier papers16
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 103 citations
- Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture SearchHouwen Peng, Hao Du, Hongyuan Yu, Qi Li et al.NeurIPS 2020 · 76 citations
- Task-Oriented Feature DistillationLinfeng Zhang, Yukang Shi, Zuoqiang Shi, Kaisheng Ma et al.NeurIPS 2020 · 74 citations
- KD-Zero: Evolving Knowledge Distiller for Any Teacher-Student PairsLujun Li, Peijie Dong, Anggeng Li, Zimian Wei et al.NeurIPS 2023 · 49 citations
- NORM: Knowledge Distillation via N-to-One Representation MatchingXiaolong Liu, Lujun Li, Chao Li, Anbang YaoICLR 2023 · 19 citations
Builds on3
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
- Knowledge Distillation via Route Constrained OptimizationXiao Jin, Baoyun Peng, Yichao Wu, Yu Liu et al.ICCV 2019 · 196 citations
- Differentiable Kernel EvolutionYu Liu, Jihao Liu, Xiaogang Wang, Ailing ZengICCV 2019 · 6 citations
Related papers
- Teacher Guided Neural Architecture Search for Face RecognitionXiaobo WangAAAI 2021 · 12 citations
- Towards Oracle Knowledge Distillation with Neural Architecture SearchMinsoo Kang, Jonghwan Mun, Bohyung HanAAAI 2020 · 48 citations
- Meta-prediction Model for Distillation-Aware NAS on Unseen DatasetsHayeon Lee, Sohyun An, Minseon Kim, Sung Ju HwangICLR 2023 · 2 citations
- Performance-Aware Mutual Knowledge Distillation for Improving Neural Architecture SearchPengtao Xie, Xuefeng DuCVPR 2022 · 16 citations
- Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language ModelsDongkuan Xu, Subhabrata Mukherjee, Xiaodong Liu, Debadeepta Dey et al.NeurIPS 2022 · 21 citations
