Progressive Network Grafting for Few-Shot Knowledge Distillation
Chengchao Shen, Xinchao Wang, Youtan Yin, Jie Song, Sihui Luo, Mingli Song
摘要
Knowledge distillation has demonstrated encouraging performances in deep model compression. Most existing approaches, however, require massive labeled data to accomplish the knowledge transfer, making the model compression a cumbersome and costly process. In this paper, we investigate the practical few-shot knowledge distillation scenario, where we assume only a few samples without human annotations are available for each category. To this end, we introduce a principled dual-stage distillation scheme tailored for fewshot data. In the first step, we graft the student blocks one by one onto the teacher, and learn the parameters of the grafted block intertwined with those of the other teacher blocks. In the second step, the trained student blocks are progressively connected and then together grafted onto the teacher network, allowing the learned student blocks to adapt themselves to each other and eventually replace the teacher network. Experiments demonstrate that our approach, with only a few unlabeled samples, achieves gratifying results on CIFAR10, CI-FAR100, and ILSVRC-2012. On CIFAR10 and CIFAR100, our performances are even on par with those of knowledge distillation schemes that utilize the full datasets. The source code is available at https://github.com/zju-vipa/NetGraft .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 被引用 103 次
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 被引用 102 次
- Mosaicking to Distill: Knowledge Distillation from Out-of-Domain DataGongfan Fang, Yifan Bao, Jie Song, Xinchao Wang 等NeurIPS 2021 · 被引用 53 次
- Meta-Aggregator: Learning to Aggregate for 1-bit Graph Neural NetworksYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song 等ICCV 2021 · 被引用 46 次
- NORM: Knowledge Distillation via N-to-One Representation MatchingXiaolong Liu, Lujun Li, Chao Li, Anbang YaoICLR 2023 · 被引用 19 次
它引用的顶会 Paper12
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang 等ICCV 2019 · 被引用 427 次
- Factorizable Graph Convolutional NetworksYiding Yang, Zunlei Feng, Mingli Song, Xinchao WangNeurIPS 2020 · 被引用 175 次
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 被引用 66 次
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge AmalgamationChengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song 等ICCV 2019 · 被引用 63 次
相关 Paper
- Few Sample Knowledge Distillation for Efficient Network CompressionTianhong Li, Jianguo Li, Zhuang Liu, Changshui ZhangCVPR 2020
- Few-Shot Class-Incremental Learning via Class-Aware Bilateral DistillationLinglan Zhao, Jing Lu, Yunlu Xu, Zhanzhan Cheng 等CVPR 2023
- Disentangle and Remerge: Interventional Knowledge Distillation for Few-Shot Object Detection from a Conditional Causal PerspectiveJiangmeng Li, Yanan Zhang, Wenwen Qiang, Lingyu Si 等AAAI 2023 · 被引用 48 次
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set SynthesisZi WangAAAI 2021 · 被引用 35 次
- All You Need in Knowledge Distillation Is a Tailored Coordinate SystemJunjie Zhou, Ke Zhu, Jianxin WuAAAI 2025 · 被引用 1 次
