Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
Hyunjune Shin, Dong-Wan Choi
摘要
Data-free knowledge distillation (DFKD) aims to distill pretrained knowledge to a student model with the help of a generator without using original data. In such data-free scenarios, achieving stable performance of DFKD is essential due to the unavailability of validation data. Unfortunately, this paper has discovered that existing DFKD methods are quite sensitive to different teacher models, occasionally showing catastrophic failures of distillation, even when using well-trained teacher models. Our observation is that the generator in DFKD is not always guaranteed to produce precise yet diverse samples using the existing representative strategy of minimizing both class-prior and adversarial losses. Through our empirical study, we focus on the fact that class-prior not only decreases the diversity of generated samples, but also cannot completely address the problem of generating unexpectedly low-quality samples depending on teacher models. In this paper, we propose the teacher-agnostic data-free knowledge distillation (TA-DFKD) method, with the goal of more robust and stable performance regardless of teacher models. Our basic idea is to assign the teacher model a lenient expert role for evaluating samples, rather than a strict supervisor that enforces its class-prior on the generator. Specifically, we design a sample selection approach that takes only clean samples verified by the teacher model without imposing restrictions on the power of generating diverse samples. Through extensive experiments, we show that our method successfully achieves both robustness and training stability across various teacher models, while outperforming the existing DFKD methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang 等ICCV 2019 · 被引用 427 次
- Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversionHongxu Yin, Pavlo Molchanov, José M. Álvarez, Zhizhong Li 等CVPR 2020
- Adaptive Data-Free QuantizationBiao Qian, Yang Wang, Richang Hong, Meng WangCVPR 2023
相关 Paper
- CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge DistillationZherui Zhang, Changwei Wang, Rongtao Xu, Wenhao Xu 等DAC 2025 · 被引用 3 次
- Sampling to Distill: Knowledge Transfer from Open-World DataYuzheng Wang, Zhaoyu Chen, Jie Zhang, Dingkang Yang 等ACM MM 2024 · 被引用 5 次
- Impartial Adversarial Distillation: Addressing Biased Data-Free Knowledge Distillation via Adaptive Constrained OptimizationDongping Liao, Xitong Gao, Chengzhong XuAAAI 2024 · 被引用 5 次
- AuG-KD: Anchor-Based Mixup Generation for Out-of-Domain Knowledge DistillationZihao Tang, Zheqi Lv, Shengyu Zhang, Yifan Zhou 等ICLR 2024 · 被引用 5 次
- Learning to Retain while Acquiring: Combating Distribution-Shift in Adversarial Data-Free Knowledge DistillationGaurav Patel, Konda Reddy Mopuri, Qiang QiuCVPR 2023
