Learning to Retain while Acquiring: Combating Distribution-Shift in Adversarial Data-Free Knowledge Distillation
Gaurav Patel, Konda Reddy Mopuri, Qiang Qiu
摘要
Data-free Knowledge Distillation (DFKD) has gained popularity recently, with the fundamental idea of carrying out knowledge transfer from a Teacher neural network to a Student neural network in the absence of training data. However, in the Adversarial DFKD framework, the student network's accuracy, suffers due to the non-stationary distribution of the pseudo-samples under multiple generator updates. To this end, at every generator update, we aim to maintain the student's performance on previously encountered examples while acquiring knowledge from samples of the current distribution. Thus, we propose a meta-learning inspired framework by treating the task of Knowledge-Acquisition (learning from newly generated samples) and Knowledge-Retention (retaining knowledge on previously met samples) as meta-train and meta-test, respectively. Hence, we dub our method as Learning to Retain while Acquiring. Moreover, we identify an implicit aligning factor between the Knowledge-Retention and Knowledge-Acquisition tasks indicating that the proposed student update strategy enforces a common gradient direction for both tasks, alleviating interference between the two objectives. Finally, we support our hypothesis by exhibiting extensive evaluation and comparison of our method with prior arts on multiple datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- DFRD: Data-Free Robustness Distillation for Heterogeneous Federated LearningKangyang Luo, Shuai Wang, Yexuan Fu, Xiang Li 等NeurIPS 2023 · 被引用 64 次
- Learning to Unlearn While Retaining: Combating Gradient Conflicts in Machine UnlearningGaurav Patel, Qiang QiuICCV 2025 · 被引用 21 次
- NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge DistillationMinh-Tuan Tran, Trung Le, Xuan-May Le, Mehrtash Harandi 等CVPR 2024 · 被引用 15 次
- Sparse Model Inversion: Efficient Inversion of Vision Transformers for Data-Free ApplicationsZixuan Hu, Yongxian Wei, Li Shen, Zhenyi Wang 等ICML 2024 · 被引用 8 次
- AuG-KD: Anchor-Based Mixup Generation for Out-of-Domain Knowledge DistillationZihao Tang, Zheqi Lv, Shengyu Zhang, Yifan Zhou 等ICLR 2024 · 被引用 5 次
它引用的顶会 Paper9
- Gradient Matching for Domain GeneralizationYuge Shi, Jeffrey Seely, Philip H. S. Torr, Siddharth Narayanaswamy 等ICLR 2022 · 被引用 358 次
- Up to 100x Faster Data-Free Knowledge DistillationGongfan Fang, Kanya Mo, Xinchao Wang, Jie Song 等AAAI 2022 · 被引用 103 次
- Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo ReplayKuluhan Binici, Shivam Aggarwal, Nam Trung Pham, Karianto Leman 等AAAI 2022 · 被引用 59 次
- Momentum Adversarial Distillation: Handling Large Distribution Shifts in Data-Free Knowledge DistillationKien Do, Hung Le, Dung Nguyen, Dang Nguyen 等NeurIPS 2022 · 被引用 48 次
- Implicit Gradient Alignment in Distributed and Federated LearningYatin Dandi, Luis Barba, Martin JaggiAAAI 2022 · 被引用 44 次
相关 Paper
- Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge DistillationHyunjune Shin, Dong-Wan ChoiAAAI 2024 · 被引用 8 次
- CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge DistillationZherui Zhang, Changwei Wang, Rongtao Xu, Wenhao Xu 等DAC 2025 · 被引用 3 次
- Sampling to Distill: Knowledge Transfer from Open-World DataYuzheng Wang, Zhaoyu Chen, Jie Zhang, Dingkang Yang 等ACM MM 2024 · 被引用 5 次
- When Data-Free Knowledge Distillation Meets Non-Transferable Teacher: Escaping Out-of-Distribution Trap is All You NeedZiming Hong, Runnan Chen, Zengmao Wang, Bo Han 等ICML 2025
- Hybrid Data-Free Knowledge DistillationJialiang Tang, Shuo Chen, Chen GongAAAI 2025 · 被引用 2 次
