Knowledge Distillation via Constrained Variational Inference
Ardavan Saeedi, Yuria Utsumi, Li Sun, Kayhan Batmanghelich, Li-Wei H. Lehman
Abstract
Knowledge distillation has been used to capture the knowledge of a teacher model and distill it into a student model with some desirable characteristics such as being smaller, more efficient, or more generalizable. In this paper, we propose a framework for distilling the knowledge of a powerful discriminative model such as a neural network into commonly used graphical models known to be more interpretable (e.g., topic models, autoregressive Hidden Markov Models). Posterior of latent variables in these graphical models (e.g., topic proportions in topic models) is often used as feature representation for predictive tasks. However, these posterior-derived features are known to have poor predictive performance compared to the features learned via purely discriminative approaches. Our framework constrains variational inference for posterior variables in graphical models with a similarity preserving constraint. This constraint distills the knowledge of the discriminative model into the graphical model by ensuring that input pairs with (dis)similar representation in the teacher model also have (dis)similar representation in the student model. By adding this constraint to the variational inference scheme, we guide the graphical model to be a reasonable density model for the data while having predictive features which are as close as possible to those of a discriminative model. To make our framework applicable to a wide range of graphical models, we build upon the Automatic Differentiation Variational Inference (ADVI), a black-box inference framework for graphical models. We demonstrate the effectiveness of our framework on two real-world tasks of disease subtyping and disease trajectory modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
- Improving Neural Topic Models using Knowledge DistillationAlexander Miserlis Hoyle, Pranav Goel, Philip ResnikEMNLP 2020 · 5 citations
Related papers
- A Unified Knowledge Distillation Framework for Deep Directed Graphical ModelsYizhuo Chen, Kaizhao Liang, Zhe Zeng, Shuochao Yao et al.CVPR 2023
- Pareto-Based Heterogeneous Knowledge Distillation for MLPs on GraphsWenrui Zhao, Yijun Tian, Zhichao Xu, Yawei Wang et al.AAAI 2026
- Knowledge Distillation with Auxiliary VariableBo Peng, Zhen Fang, Guangquan Zhang, Jie LuICML 2024 · 7 citations
- Do Topological Characteristics Help in Knowledge Distillation?Jungeun Kim, Junwon You, Dongjin Lee, Ha Young Kim et al.ICML 2024 · 11 citations
- Predictive variational inference: Learn the predictively optimal posterior distributionJinlin Lai, Antonio Linero, Yuling YaoICML 2026
