Progressive Network Grafting for Few-Shot Knowledge Distillation
Chengchao Shen, Xinchao Wang, Youtan Yin, Jie Song, Sihui Luo, Mingli Song
Abstract
Knowledge distillation has demonstrated encouraging performances in deep model compression. Most existing approaches, however, require massive labeled data to accomplish the knowledge transfer, making the model compression a cumbersome and costly process. In this paper, we investigate the practical few-shot knowledge distillation scenario, where we assume only a few samples without human annotations are available for each category. To this end, we introduce a principled dual-stage distillation scheme tailored for fewshot data. In the first step, we graft the student blocks one by one onto the teacher, and learn the parameters of the grafted block intertwined with those of the other teacher blocks. In the second step, the trained student blocks are progressively connected and then together grafted onto the teacher network, allowing the learned student blocks to adapt themselves to each other and eventually replace the teacher network. Experiments demonstrate that our approach, with only a few unlabeled samples, achieves gratifying results on CIFAR10, CI-FAR100, and ILSVRC-2012. On CIFAR10 and CIFAR100, our performances are even on par with those of knowledge distillation schemes that utilize the full datasets. The source code is available at https://github.com/zju-vipa/NetGraft .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6279a3a6-63d3-4213-b8c4-c463773db91fCited by top-tier papers18
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 103 citations
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 102 citations
- Mosaicking to Distill: Knowledge Distillation from Out-of-Domain DataGongfan Fang, Yifan Bao, Jie Song, Xinchao Wang et al.NeurIPS 2021 · 53 citations
- Meta-Aggregator: Learning to Aggregate for 1-bit Graph Neural NetworksYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song et al.ICCV 2021 · 46 citations
- NORM: Knowledge Distillation via N-to-One Representation MatchingXiaolong Liu, Lujun Li, Chao Li, Anbang YaoICLR 2023 · 19 citations
Builds on12
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang et al.ICCV 2019 · 427 citations
- Factorizable Graph Convolutional NetworksYiding Yang, Zunlei Feng, Mingli Song, Xinchao WangNeurIPS 2020 · 175 citations
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 66 citations
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge AmalgamationChengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song et al.ICCV 2019 · 63 citations
Related papers
- Few Sample Knowledge Distillation for Efficient Network CompressionTianhong Li, Jianguo Li, Zhuang Liu, Changshui ZhangCVPR 2020
- Few-Shot Class-Incremental Learning via Class-Aware Bilateral DistillationLinglan Zhao, Jing Lu, Yunlu Xu, Zhanzhan Cheng et al.CVPR 2023
- Disentangle and Remerge: Interventional Knowledge Distillation for Few-Shot Object Detection from a Conditional Causal PerspectiveJiangmeng Li, Yanan Zhang, Wenwen Qiang, Lingyu Si et al.AAAI 2023 · 48 citations
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set SynthesisZi WangAAAI 2021 · 35 citations
- All You Need in Knowledge Distillation Is a Tailored Coordinate SystemJunjie Zhou, Ke Zhu, Jianxin WuAAAI 2025 · 1 citation
