Up to 100x Faster Data-Free Knowledge Distillation
Gongfan Fang, Kanya Mo, Xinchao Wang, Jie Song, Shitao Bei, Haofei Zhang, Mingli Song
Abstract
Data-free knowledge distillation (DFKD) has recently been attracting increasing attention from research communities, attributed to its capability to compress a model only using synthetic data. Despite the encouraging results achieved, stateof-the-art DFKD methods still suffer from the inefficiency of data synthesis, making the data-free training process extremely time-consuming and thus inapplicable for large-scale tasks. In this work, we introduce an efficacious scheme, termed as FastDFKD, that allows us to accelerate DFKD by a factor of orders of magnitude. At the heart of our approach is a novel strategy to reuse the shared common features in training data so as to synthesize different data instances. Unlike prior methods that optimize a set of data independently, we propose to learn a meta-synthesizer that seeks common features as the initialization for the fast data synthesis. As a result, FastDFKD achieves data synthesis within only a few steps, significantly enhancing the efficiency of data-free training. Experiments over CIFAR, NYUv2, and ImageNet demonstrate that the proposed FastDFKD achieves 10× and even 100× acceleration while preserving performances on par with state of the art. Code is available at https://github.com/zju-vipa/Fast-Datafree . 2020; Fang et al. 2021a). The conventional setup of KD requires possessing the original training data as input so as to train the student. Unfortunately, due to confidential or copyright reasons, in many cases the original data cannot be released and only the pre-trained models are available to users (Kolesnikov et al. 2020; Shen et al. 2019; Ye et al. 2019) , which, in turn, imposes a major obstacle towards applying KD for a broader domain. To remedy this issue, data-free knowledge distillation (DFKD) approaches have been proposed by assuming that no access to training data is available at all (Lopes, Fenu, and
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88989331-a322-46e4-a11b-3c51d892e4fcCited by top-tier papers40
- TARGET: Federated Class-Continual Learning via Exemplar-Free DistillationJie Zhang, Chen Chen, Weiming Zhuang, Lingjuan LyuICCV 2023 · 109 citations
- Momentum Adversarial Distillation: Handling Large Distribution Shifts in Data-Free Knowledge DistillationKien Do, Hung Le, Dung Nguyen, Dang Nguyen et al.NeurIPS 2022 · 48 citations
- Lion: Adversarial Distillation of Proprietary Large Language ModelsYuxin Jiang, Chunkit Chan, Mingyang Chen, Wei WangEMNLP 2023 · 20 citations
- Are Large Kernels Better Teachers than Transformers for ConvNets?Tianjin Huang, Lu Yin, Zhenyu Zhang, Li Shen et al.ICML 2023 · 18 citations
- Revisiting Data-Free Knowledge Distillation with Poisoned TeachersJunyuan Hong, Yi Zeng, Shuyang Yu, Lingjuan Lyu et al.ICML 2023 · 16 citations
Builds on6
- Data-Free Knowledge Distillation for Heterogeneous Federated LearningZhuangdi Zhu, Junyuan Hong, Jiayu ZhouICML 2021 · 957 citations
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang et al.ICCV 2019 · 427 citations
- Adversarial Self-Supervised Data-Free Distillation for Text ClassificationXinyin Ma, Yongliang Shen, Gongfan Fang, Chen Chen et al.EMNLP 2020 · 18 citations
- Distilling Knowledge From Graph Convolutional NetworksYiding Yang, Jiayan Qiu, Mingli Song, Dacheng Tao et al.CVPR 2020
- Learning Student Networks in the WildHanting Chen, Tianyu Guo, Chang Xu, Wenshuo Li et al.CVPR 2021
Related papers
- Small Scale Data-Free Knowledge DistillationHe Liu, Yikai Wang, Huaping Liu, Fuchun Sun et al.CVPR 2024
- Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion AugmentationMuquan Li, Dongyang Zhang, Tao He, Xiurui Xie et al.ACM MM 2024 · 5 citations
- Sampling to Distill: Knowledge Transfer from Open-World DataYuzheng Wang, Zhaoyu Chen, Jie Zhang, Dingkang Yang et al.ACM MM 2024 · 5 citations
- IFHE: Intermediate-Feature Heterogeneity Enhancement for Image Synthesis in Data-Free Knowledge DistillationYi Chen, Ning Liu, Ao Ren, Tao Yang et al.DAC 2023 · 1 citation
- CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge DistillationZherui Zhang, Changwei Wang, Rongtao Xu, Wenhao Xu et al.DAC 2025 · 3 citations
