DisWOT: Student Architecture Search for Distillation WithOut Training
Peijie Dong, Lujun Li, Zimian Wei
摘要
Knowledge distillation (KD) is an effective training strategy to improve the lightweight student models under the guidance of cumbersome teachers. However, the large architecture difference across the teacher-student pairs limits the distillation gains. In contrast to previous adaptive distillation methods to reduce the teacher-student gap, we explore a novel training-free framework to search for the best student architectures for a given teacher. Our work first empirically show that the optimal model under vanilla training cannot be the winner in distillation. Secondly, we find that the similarity of feature semantics and sample relations between random-initialized teacherstudent networks have good correlations with final distillation performances. Thus, we efficiently measure similarity matrixs conditioned on the semantic activation maps to select the optimal student via an evolutionary algorithm without any training. In this way, our student architecture search for Distillation WithOut Training (DisWOT) significantly improves the performance of the model in the distillation stage with at least 180× training acceleration. Additionally, we extend similarity metrics in DisWOT as new distillers and KD-based zero-proxies. Our experiments on CIFAR, ImageNet and NAS-Bench-201 demonstrate that our technique achieves state-of-the-art results on different search spaces. Our project and code are available at https://lilujunai.github.io/DisWOT-CVPR2023/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language ModelsPeijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu 等ICML 2024 · 被引用 64 次
- Automated Knowledge Distillation via Monte Carlo Tree SearchLujun Li, Peijie Dong, Zimian Wei, Ya YangICCV 2023 · 被引用 54 次
- Discovering Sparsity Allocation for Layer-wise Pruning of Large Language ModelsLujun Li, Peijie Dong, Zhenheng Tang, Xiang Liu 等NeurIPS 2024 · 被引用 51 次
- EMQ: Evolving Training-free Proxies for Automated Mixed Precision QuantizationPeijie Dong, Lujun Li, Zimian Wei, Xin Niu 等ICCV 2023 · 被引用 51 次
- KD-Zero: Evolving Knowledge Distiller for Any Teacher-Student PairsLujun Li, Peijie Dong, Anggeng Li, Zimian Wei 等NeurIPS 2023 · 被引用 49 次
它引用的顶会 Paper29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
相关 Paper
- Towards Oracle Knowledge Distillation with Neural Architecture SearchMinsoo Kang, Jonghwan Mun, Bohyung HanAAAI 2020 · 被引用 48 次
- UniADS: Universal Architecture-Distiller Search for Distillation GapLiming Lu, Zhenghan Chen, Xiaoyu Lu, Yihang Rao 等AAAI 2024 · 被引用 19 次
- Adaptive Dual Guidance Knowledge DistillationTong Li, Long Liu, Kang Liu, Xin Wang 等AAAI 2025 · 被引用 1 次
- Search to Distill: Pearls Are Everywhere but Not the EyesYu Liu, Xuhui Jia, Mingxing Tan, Raviteja Vemulapalli 等CVPR 2020
- Knowledge Diffusion for DistillationTao Huang, Yuan Zhang, Mingkai Zheng, Shan You 等NeurIPS 2023 · 被引用 125 次
