Few Sample Knowledge Distillation for Efficient Network Compression
Tianhong Li, Jianguo Li, Zhuang Liu, Changshui Zhang
摘要
Deep neural network compression techniques such as pruning and weight tensor decomposition usually require fine-tuning to recover the prediction accuracy when the compression ratio is high. However, conventional finetuning suffers from the requirement of a large training set and the time-consuming training procedure. This paper proposes a novel solution for knowledge distillation from label-free few samples to realize both data efficiency and training/processing efficiency. We treat the original network as "teacher-net" and the compressed network as "studentnet". A 1×1 convolution layer is added at the end of each layer block of the student-net, and we fit the block-level outputs of the student-net to the teacher-net by estimating the parameters of the added layers. We prove that the added layer can be merged without adding extra parameters and computation cost during inference. Experiments on multiple datasets and network architectures verify the method's effectiveness on student-nets obtained by various network pruning and weight decomposition methods. Our method can recover student-net's accuracy to the same level as conventional fine-tuning methods in minutes while using only 1% label-free data of the full training data. * This work was done when Tianhong Li was intern at Intel Labs under supervised by Jianguo Li.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot 等ICLR 2020 · 被引用 244 次
- Revisiting Random Channel Pruning for Neural Network CompressionYawei Li, Kamil Adamczewski, Wen Li, Shuhang Gu 等CVPR 2022 · 被引用 114 次
- Binocular Mutual Learning for Improving Few-shot ClassificationZiqi Zhou, Xi Qiu, Jiangtao Xie, Jianan Wu 等ICCV 2021 · 被引用 101 次
- Mistify: Automating DNN Model Porting for On-Device Inference at the EdgePeizhen Guo, Bo Hu, Wenjun HuNSDI 2021 · 被引用 69 次
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge AmalgamationChengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song 等ICCV 2019 · 被引用 63 次
它引用的顶会 Paper2
相关 Paper
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 被引用 66 次
- Progressive Network Grafting for Few-Shot Knowledge DistillationChengchao Shen, Xinchao Wang, Youtan Yin, Jie Song 等AAAI 2021 · 被引用 55 次
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set SynthesisZi WangAAAI 2021 · 被引用 35 次
- Beyond Student: An Asymmetric Network for Neural Network InheritanceYiyun Zhou, Jingwei Shi, Mingjing Xu, Zhonghua Jiang 等ICLR 2026 · 被引用 1 次
- Hybrid Data-Free Knowledge DistillationJialiang Tang, Shuo Chen, Chen GongAAAI 2025 · 被引用 2 次
