Knowledge Distillation via Constrained Variational Inference
Ardavan Saeedi, Yuria Utsumi, Li Sun, Kayhan Batmanghelich, Li-Wei H. Lehman
摘要
Knowledge distillation has been used to capture the knowledge of a teacher model and distill it into a student model with some desirable characteristics such as being smaller, more efficient, or more generalizable. In this paper, we propose a framework for distilling the knowledge of a powerful discriminative model such as a neural network into commonly used graphical models known to be more interpretable (e.g., topic models, autoregressive Hidden Markov Models). Posterior of latent variables in these graphical models (e.g., topic proportions in topic models) is often used as feature representation for predictive tasks. However, these posterior-derived features are known to have poor predictive performance compared to the features learned via purely discriminative approaches. Our framework constrains variational inference for posterior variables in graphical models with a similarity preserving constraint. This constraint distills the knowledge of the discriminative model into the graphical model by ensuring that input pairs with (dis)similar representation in the teacher model also have (dis)similar representation in the student model. By adding this constraint to the variational inference scheme, we guide the graphical model to be a reasonable density model for the data while having predictive features which are as close as possible to those of a discriminative model. To make our framework applicable to a wide range of graphical models, we build upon the Automatic Differentiation Variational Inference (ADVI), a black-box inference framework for graphical models. We demonstrate the effectiveness of our framework on two real-world tasks of disease subtyping and disease trajectory modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 被引用 59 次
- Improving Neural Topic Models using Knowledge DistillationAlexander Miserlis Hoyle, Pranav Goel, Philip ResnikEMNLP 2020 · 被引用 5 次
相关 Paper
- A Unified Knowledge Distillation Framework for Deep Directed Graphical ModelsYizhuo Chen, Kaizhao Liang, Zhe Zeng, Shuochao Yao 等CVPR 2023
- Pareto-Based Heterogeneous Knowledge Distillation for MLPs on GraphsWenrui Zhao, Yijun Tian, Zhichao Xu, Yawei Wang 等AAAI 2026
- Knowledge Distillation with Auxiliary VariableBo Peng, Zhen Fang, Guangquan Zhang, Jie LuICML 2024 · 被引用 7 次
- Do Topological Characteristics Help in Knowledge Distillation?Jungeun Kim, Junwon You, Dongjin Lee, Ha Young Kim 等ICML 2024 · 被引用 11 次
- Predictive variational inference: Learn the predictively optimal posterior distributionJinlin Lai, Antonio Linero, Yuling YaoICML 2026
