Improving Topic Modeling by Distilling Soft Labels from Language Models
Raymond Li, Amirhossein Abaskohi, Chuyuan Li, Gabriel Murray, Giuseppe Carenini
Abstract
Traditional neural topic models are typically optimized by reconstructing the document's Bag-of-Words (BoW) representations, overlooking contextual information and struggling with data sparsity. In this work, we introduce a novel topic model training framework by Distilling Soft Labels (DSL) from Language Models (LMs). To construct the contextually enriched reconstruction signals, we project the next token probabilities, conditioned on a specialized prompt, onto a pre-defined vocabulary, and train the topic models to reconstruct the soft labels using the LM hidden states. This produces higher-quality topics that are more closely aligned with the underlying thematic structure of the corpus. Extensive experiments demonstrate that DSL achieves substantial improvements in topic coherence and assignment accuracy over existing baselines. Additionally, we also introduce a retrieval-based metric, which shows that our approach significantly outperforms existing methods in identifying semantically similar documents, highlighting its effectiveness for retrieval-oriented applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c1a365d-2e03-4475-bf83-402cda5bd792Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceAlexander Miserlis Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov et al.NeurIPS 2021 · 220 citations
- Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context LearningXinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers et al.NeurIPS 2023 · 206 citations
Related papers
- Adaptive Pseudo-Labeling via Word Coherence for Topic ModelingBohan Yoon, Hyejin JangKDD 2026
- Improving Neural Topic Models using Knowledge DistillationAlexander Miserlis Hoyle, Pranav Goel, Philip ResnikEMNLP 2020 · 5 citations
- TAN-NTM: Topic Attention Networks for Neural Topic ModelingMadhur Panwar, Shashank Shailabh, Milan Aggarwal, Balaji KrishnamurthyACL 2021
- Contrastive Learning for Neural Topic ModelThong Nguyen, Anh Tuan LuuNeurIPS 2021 · 82 citations
- Explainable and Discourse Topic-aware Neural Language UnderstandingYatin Chaudhary, Hinrich Schütze, Pankaj GuptaICML 2020 · 7 citations
