Small Scale Data-Free Knowledge Distillation
He Liu, Yikai Wang, Huaping Liu, Fuchun Sun, Anbang Yao
摘要
Data-free knowledge distillation is able to utilize the knowledge learned by a large teacher network to augment the training of a smaller student network without accessing the original training data, avoiding privacy, security, and proprietary risks in real applications. In this line of research, existing methods typically follow an inversion-anddistillation paradigm in which a generative adversarial network on-the-fly trained with the guidance of the pre-trained teacher network is used to synthesize a large-scale sample set for knowledge distillation. In this paper, we reexamine this common data-free knowledge distillation paradigm, showing that there is considerable room to improve the overall training efficiency through a lens of "small-scale inverted data for knowledge distillation". In light of three empirical observations indicating the importance of how to balance class distributions in terms of synthetic sample diversity and difficulty during both data inversion and distillation processes, we propose Small Scale Data-free Knowledge Distillation (SSD-KD). In formulation, SSD-KD introduces a modulating function to balance synthetic samples and a priority sampling function to select proper samples, facilitated by a dynamic replay buffer and a reinforcement learning strategy. As a result, SSD-KD can perform distillation training conditioned on an extremely small scale of synthetic samples (e.g., 10× less than the original training data scale), making the overall training efficiency one or two orders of magnitude faster than many mainstream methods while retaining superior or competitive model performance, as demonstrated on popular image classification and semantic segmentation benchmarks. The code is available at https://github.com/OSVAI/SSD-KD . * Equal contribution. † Corresponding author. This work was done when He Liu was an intern at Intel Labs China, supervised by Anbang Yao who conceived the project. planemobile bird cat deer dog frog horse ship truck Category 50 55 60 65 70 75 80 85 90 95 Accuracy (%) CIFAR-10 (T:ResNet34, S:ResNet18) Vanilla KD DeepInv SSD-KD planemobile bird cat deer dog frog horse ship truck Category 50 55 60 65 70 75 80 85 90 95 Accuracy (%) CIFAR-10 (T:VGG11, S:ResNet18) Vanilla KD DeepInv SSD-KD planemobile bird cat deer dog frog horse ship truck Category 50 55 60 65 70 75 80 85 90 95 Accuracy (%)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge DistillationZherui Zhang, Changwei Wang, Rongtao Xu, Wenhao Xu 等DAC 2025 · 被引用 3 次
- RUAGO: Effective and Practical Retain-Free Unlearning via Adversarial Attack and OOD GeneratorSangyong Lee, Sangjun Chung, Simon S. WooNeurIPS 2025 · 被引用 3 次
- Scaling Dense Event-Stream Pretraining from Visual Foundation ModelsZhiwen Chen, Junhui Hou, Zhiyu Zhu, Jinjian Wu 等CVPR 2026 · 被引用 2 次
- TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRAChanJoo Jung, Jaehyung KimICLR 2026 · 被引用 2 次
- Coupling the Generator with Teacher for Effective Data-Free Knowledge DistillationXu Chen, Yang Li, Yahong Han, Guangquan Xu 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper10
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang 等ICCV 2019 · 被引用 427 次
- Always Be Dreaming: A New Approach for Data-Free Class-Incremental LearningJames Seale Smith, Yen-Chang Hsu, Jonathan C. Balloch, Yilin Shen 等ICCV 2021 · 被引用 208 次
- Up to 100x Faster Data-Free Knowledge DistillationGongfan Fang, Kanya Mo, Xinchao Wang, Jie Song 等AAAI 2022 · 被引用 103 次
- MixMix: All You Need for Data-Free Compression Are Feature and Data MixingYuhang Li, Feng Zhu, Ruihao Gong, Mingzhu Shen 等ICCV 2021 · 被引用 52 次
相关 Paper
- Hybrid Data-Free Knowledge DistillationJialiang Tang, Shuo Chen, Chen GongAAAI 2025 · 被引用 2 次
- Sampling to Distill: Knowledge Transfer from Open-World DataYuzheng Wang, Zhaoyu Chen, Jie Zhang, Dingkang Yang 等ACM MM 2024 · 被引用 5 次
- Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo ReplayKuluhan Binici, Shivam Aggarwal, Nam Trung Pham, Karianto Leman 等AAAI 2022 · 被引用 59 次
- AuG-KD: Anchor-Based Mixup Generation for Out-of-Domain Knowledge DistillationZihao Tang, Zheqi Lv, Shengyu Zhang, Yifan Zhou 等ICLR 2024 · 被引用 5 次
- Momentum Adversarial Distillation: Handling Large Distribution Shifts in Data-Free Knowledge DistillationKien Do, Hung Le, Dung Nguyen, Dang Nguyen 等NeurIPS 2022 · 被引用 48 次
