Improving Neural Topic Models using Knowledge Distillation
Alexander Miserlis Hoyle, Pranav Goel, Philip Resnik
Abstract
Topic models are often used to identify humaninterpretable topics to help make sense of large document collections. We use knowledge distillation to combine the best attributes of probabilistic topic models and pretrained transformers. Our modular method can be straightforwardly applied with any neural topic model to improve topic quality, which we demonstrate using two models having disparate architectures, obtaining state-of-the-art topic coherence. We show that our adaptable framework not only improves performance in the aggregate over all estimated topics, as is commonly reported, but also in head-to-head comparisons of aligned topics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90d192df-8156-4290-888c-aefdc26ec39aCited by top-tier papers12
- Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceAlexander Miserlis Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov et al.NeurIPS 2021 · 220 citations
- Contrastive Learning for Neural Topic ModelThong Nguyen, Anh Tuan LuuNeurIPS 2021 · 82 citations
- Topic Discovery via Latent Space Clustering of Pretrained Language Model RepresentationsYu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang et al.WWW 2022 · 73 citations
- Knowledge-Aware Bayesian Deep Topic ModelDongsheng Wang, Yishi Xu, Miaoge Li, Zhibin Duan et al.NeurIPS 2022 · 19 citations
- Pre-training and Fine-tuning Neural Topic Model: A Simple yet Effective Approach to Incorporating External KnowledgeLinhai Zhang, Xuemeng Hu, Boyu Wang, Deyu Zhou et al.ACL 2022 · 14 citations
Builds on3
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama et al.ACL 2020 · 37 citations
- Creating Something From Nothing: Unsupervised Knowledge Distillation for Cross-Modal HashingHengtong Hu, Lingxi Xie, Richang Hong, Qi TianCVPR 2020
Related papers
- Friendly Topic Assistant for Transformer Based Abstractive SummarizationZhengjue Wang, Zhibin Duan, Hao Zhang, Chaojie Wang et al.EMNLP 2020 · 48 citations
- Disentangling Transformer Language Models as Superposed Topic ModelsJia Peng Lim, Hady W. LauwEMNLP 2023 · 2 citations
- Improving Topic Modeling by Distilling Soft Labels from Language ModelsRaymond Li, Amirhossein Abaskohi, Chuyuan Li, Gabriel Murray et al.ICML 2026
- A Discrete Variational Recurrent Topic Model without the Reparametrization TrickMehdi Rezaee, Francis FerraroNeurIPS 2020 · 31 citations
- Neural Topic Modeling with Large Language Models in the LoopXiaohao Yang, He Zhao, Weijie Xu, Yuanyuan Qi et al.ACL 2025 · 13 citations
