A Unified Knowledge Distillation Framework for Deep Directed Graphical Models
Yizhuo Chen, Kaizhao Liang, Zhe Zeng, Shuochao Yao, Huajie Shao
Abstract
Knowledge distillation (KD) is a technique that transfers the knowledge from a large teacher network to a small student network. It has been widely applied to many different tasks, such as model compression and federated learning. However, existing KD methods fail to generalize to general deep directed graphical models (DGMs) with arbitrary layers of random variables. We refer by deep DGMs to DGMs whose conditional distributions are parameterized by deep neural networks. In this work, we propose a novel unified knowledge distillation framework for deep DGMs on various applications. Specifically, we leverage the reparameterization trick to hide the intermediate latent variables, resulting in a compact DGM. Then we develop a surrogate distillation loss to reduce error accumulation through multiple layers of random variables. Moreover, we present the connections between our method and some existing knowledge distillation approaches. The proposed framework is evaluated on four applications: data-free hierarchical variational autoencoder (VAE) compression, data-free variational recurrent neural networks (VRNN) compression, data-free Helmholtz Machine (HM) compression, and VAE continual learning. The results show that our distillation method outperforms the baselines in data-free model compression tasks. We further demonstrate that our method significantly improves the performance of KD-based continual learning for data generation. Our source code is available at https: //github.com/YizhuoChen99/KD4DGM-CVPR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11de1834-54ba-4a48-aebb-6d5b67602b9dBuilds on7
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 1,615 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Lifelong GAN: Continual Learning for Conditional Image GenerationMengyao Zhai, Lei Chen, Frederick Tung, Jiawei He et al.ICCV 2019 · 204 citations
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 103 citations
- Autoregressive Knowledge Distillation through Imitation LearningAlexander Lin, Jeremy Wohlwend, Howard Chen, Tao LeiEMNLP 2020 · 18 citations
Related papers
- Data-Free Knowledge Distillation with Soft Targeted Transfer Set SynthesisZi WangAAAI 2021 · 35 citations
- Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo ReplayKuluhan Binici, Shivam Aggarwal, Nam Trung Pham, Karianto Leman et al.AAAI 2022 · 59 citations
- Data-Free Ensemble Knowledge Distillation for Privacy-conscious Multimedia Model CompressionZhiwei Hao, Yong Luo, Han Hu, Jianping An et al.ACM MM 2021 · 11 citations
- FreeKD: Free-direction Knowledge Distillation for Graph Neural NetworksKaituo Feng, Changsheng Li, Ye Yuan, Guoren WangKDD 2022 · 28 citations
- Learning to Retain while Acquiring: Combating Distribution-Shift in Adversarial Data-Free Knowledge DistillationGaurav Patel, Konda Reddy Mopuri, Qiang QiuCVPR 2023
