Topic Discovery via Latent Space Clustering of Pretrained Language Model Representations
Yu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang, Jiawei Han
Abstract
Topic models have been the prominent tools for automatic topic discovery from text corpora. Despite their effectiveness, topic models suffer from several limitations including the inability of modeling word ordering information in documents, the difficulty of incorporating external linguistic knowledge, and the lack of both accurate and efficient inference methods for approximating the intractable posterior. Recently, pretrained language models (PLMs) have brought astonishing performance improvements to a wide variety of tasks due to their superior representations of text. Interestingly, there have not been standard approaches to deploy PLMs for topic discovery as better alternatives to topic models. In this paper, we begin by analyzing the challenges of using PLM representations for topic discovery, and then propose a joint latent space learning and clustering framework built upon PLM embeddings. In the latent space, topic-word and document-topic distributions are jointly modeled so that the discovered topics can be interpreted by coherent and distinctive terms and meanwhile serve as meaningful summaries of the documents. Our model effectively leverages the strong representation power and superb linguistic features brought by PLMs for topic discovery, and is conceptually simpler than topic models. On two benchmark datasets in different domains, our model generates significantly more coherent and diverse topics than strong topic models, and offers better topic-wise document representations, based on both automatic and human evaluations. 1 CCS CONCEPTS • Information systems → Clustering; Document topic models; • Computing methodologies → Natural language processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ceacd582-14f8-4bed-8502-2e7d82eb84a7Cited by top-tier papers11
- How Do Transformers Learn Topic Structure: Towards a Mechanistic UnderstandingYuchen Li, Yuanzhi Li, Andrej RisteskiICML 2023 · 87 citations
- Heterformer: Transformer-based Deep Node Representation Learning on Heterogeneous Text-Rich NetworksBowen Jin, Yu Zhang, Qi Zhu, Jiawei HanKDD 2023 · 27 citations
- Specious Sites: Tracking the Spread and Sway of Spurious News Stories at ScaleHans W. A. Hanley, Deepak Kumar, Zakir DurumericS&P 2024 · 18 citations
- Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document GenerationYoungjoon Yoo, Jongwon ChoiAAAI 2024 · 8 citations
- Emergence of Hierarchical Emotion Organization in Large Language ModelsMaya Okawa, Bo Zhao, Eric Bigelow, Rose Yu et al.ICML 2026 · 4 citations
Builds on12
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model PretrainingYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary et al.NeurIPS 2021 · 231 citations
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong et al.EMNLP 2020 · 203 citations
- Discriminative Topic Mining via Category-Name Guided Text EmbeddingYu Meng, Jiaxin Huang, Guangyuan Wang, Zihan Wang et al.WWW 2020 · 80 citations
Related papers
- Neural Topic Modeling with Large Language Models in the LoopXiaohao Yang, He Zhao, Weijie Xu, Yuanyuan Qi et al.ACL 2025 · 13 citations
- Representing Mixtures of Word Embeddings with Mixtures of Topic EmbeddingsDongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng et al.ICLR 2022 · 56 citations
- Explainable and Discourse Topic-aware Neural Language UnderstandingYatin Chaudhary, Hinrich Schütze, Pankaj GuptaICML 2020 · 7 citations
- Beyond prompting: Making Pre-trained Language Models Better Zero-shot Learners by Clustering RepresentationsYu Fei, Zhao Meng, Ping Nie, Roger Wattenhofer et al.EMNLP 2022 · 13 citations
- Pre-training and Fine-tuning Neural Topic Model: A Simple yet Effective Approach to Incorporating External KnowledgeLinhai Zhang, Xuemeng Hu, Boyu Wang, Deyu Zhou et al.ACL 2022 · 14 citations
