Momentum Adversarial Distillation: Handling Large Distribution Shifts in Data-Free Knowledge Distillation
Kien Do, Hung Le, Dung Nguyen, Dang Nguyen, Haripriya Harikumar, Truyen Tran, Santu Rana, Svetha Venkatesh
Abstract
Data-free Knowledge Distillation (DFKD) has attracted attention recently thanks to its appealing capability of transferring knowledge from a teacher network to a student network without using training data. The main idea is to use a generator to synthesize data for training the student. As the generator gets updated, the distribution of synthetic data will change. Such distribution shift could be large if the generator and the student are trained adversarially, causing the student to forget the knowledge it acquired at previous steps. To alleviate this problem, we propose a simple yet effective method called Momentum Adversarial Distillation (MAD) which maintains an exponential moving average (EMA) copy of the generator and uses synthetic samples from both the generator and the EMA generator to train the student. Since the EMA generator can be considered as an ensemble of the generator's old versions and often undergoes a smaller change in updates compared to the generator, training on its synthetic samples can help the student recall the past knowledge and prevent the student from adapting too quickly to new updates of the generator. Our experiments on six benchmark datasets including big datasets like ImageNet and Places365 demonstrate the superior performance of MAD over competing methods for handling the large distribution shift problem. Our method also compares favorably to existing DFKD methods and even achieves state-of-the-art results in some cases. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad3dac71-8bab-4bca-bfc8-2bc63d9f6d17Cited by top-tier papers18
- DFRD: Data-Free Robustness Distillation for Heterogeneous Federated LearningKangyang Luo, Shuai Wang, Yexuan Fu, Xiang Li et al.NeurIPS 2023 · 64 citations
- Distribution Shift Matters for Knowledge Distillation with Webly Collected ImagesJialiang Tang, Shuo Chen, Gang Niu, Masashi Sugiyama et al.ICCV 2023 · 21 citations
- NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge DistillationMinh-Tuan Tran, Trung Le, Xuan-May Le, Mehrtash Harandi et al.CVPR 2024 · 15 citations
- Data-Free Hard-Label Robustness Stealing AttackXiaojian Yuan, Kejiang Chen, Wen Huang, Jie Zhang et al.AAAI 2024 · 11 citations
- De-Confounded Data-Free Knowledge Distillation for Handling Distribution ShiftsYuzheng Wang, Dingkang Yang, Zhaoyu Chen, Yang Liu et al.CVPR 2024 · 10 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang et al.ICCV 2019 · 427 citations
- Always Be Dreaming: A New Approach for Data-Free Class-Incremental LearningJames Seale Smith, Yen-Chang Hsu, Jonathan C. Balloch, Yilin Shen et al.ICCV 2021 · 208 citations
Related papers
- Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo ReplayKuluhan Binici, Shivam Aggarwal, Nam Trung Pham, Karianto Leman et al.AAAI 2022 · 59 citations
- Learning to Retain while Acquiring: Combating Distribution-Shift in Adversarial Data-Free Knowledge DistillationGaurav Patel, Konda Reddy Mopuri, Qiang QiuCVPR 2023
- Hybrid Data-Free Knowledge DistillationJialiang Tang, Shuo Chen, Chen GongAAAI 2025 · 2 citations
- Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge DistillationHyunjune Shin, Dong-Wan ChoiAAAI 2024 · 8 citations
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set SynthesisZi WangAAAI 2021 · 35 citations
