Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay
Kuluhan Binici, Shivam Aggarwal, Nam Trung Pham, Karianto Leman, Tulika Mitra
摘要
Data-Free Knowledge Distillation (KD) allows knowledge transfer from a trained neural network (teacher) to a more compact one (student) in the absence of original training data. Existing works use a validation set to monitor the accuracy of the student over real data and report the highest performance throughout the entire process. However, validation data may not be available at distillation time either, making it infeasible to record the student snapshot that achieved the peak accuracy. Therefore, a practical data-free KD method should be robust and ideally provide monotonically increasing student accuracy during distillation. This is challenging because the student experiences knowledge degradation due to the distribution shift of the synthetic data. A straightforward approach to overcome this issue is to store and rehearse the generated samples periodically, which increases the memory footprint and creates privacy concerns. We propose to model the distribution of the previously observed synthetic samples with a generative network. In particular, we design a Variational Autoencoder (VAE) with a training objective that is customized to learn the synthetic data representations optimally. The student is rehearsed by the generative pseudo replay technique, with samples produced by the VAE. Hence knowledge degradation can be prevented without storing any samples. Experiments on image classification benchmarks show that our method optimizes the expected value of the distilled model accuracy while eliminating the large memory overhead incurred by the sample-storing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- DFRD: Data-Free Robustness Distillation for Heterogeneous Federated LearningKangyang Luo, Shuai Wang, Yexuan Fu, Xiang Li 等NeurIPS 2023 · 被引用 64 次
- Momentum Adversarial Distillation: Handling Large Distribution Shifts in Data-Free Knowledge DistillationKien Do, Hung Le, Dung Nguyen, Dang Nguyen 等NeurIPS 2022 · 被引用 48 次
- Distribution Shift Matters for Knowledge Distillation with Webly Collected ImagesJialiang Tang, Shuo Chen, Gang Niu, Masashi Sugiyama 等ICCV 2023 · 被引用 21 次
- NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge DistillationMinh-Tuan Tran, Trung Le, Xuan-May Le, Mehrtash Harandi 等CVPR 2024 · 被引用 15 次
- MAP: MAsk-Pruning for Source-Free Model Intellectual Property ProtectionBoyang Peng, Sanqing Qu, Yong Wu, Tianpei Zou 等CVPR 2024 · 被引用 6 次
它引用的顶会 Paper7
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang 等ICCV 2019 · 被引用 427 次
- Membership Leakage in Label-Only ExposuresZheng Li, Yang ZhangCCS 2021 · 被引用 185 次
- DeGAN: Data-Enriching GAN for Retrieving Representative Samples from a Trained ClassifierSravanti Addepalli, Gaurav Kumar Nayak, Anirban Chakraborty, Venkatesh Babu RadhakrishnanAAAI 2020 · 被引用 40 次
- Adversarial Self-Supervised Data-Free Distillation for Text ClassificationXinyin Ma, Yongliang Shen, Gongfan Fang, Chen Chen 等EMNLP 2020 · 被引用 18 次
- Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversionHongxu Yin, Pavlo Molchanov, José M. Álvarez, Zhizhong Li 等CVPR 2020
相关 Paper
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set SynthesisZi WangAAAI 2021 · 被引用 35 次
- Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge DistillationHyunjune Shin, Dong-Wan ChoiAAAI 2024 · 被引用 8 次
- Hybrid Data-Free Knowledge DistillationJialiang Tang, Shuo Chen, Chen GongAAAI 2025 · 被引用 2 次
- Learning to Retain while Acquiring: Combating Distribution-Shift in Adversarial Data-Free Knowledge DistillationGaurav Patel, Konda Reddy Mopuri, Qiang QiuCVPR 2023
- CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge DistillationZherui Zhang, Changwei Wang, Rongtao Xu, Wenhao Xu 等DAC 2025 · 被引用 3 次
