Topic Discovery via Latent Space Clustering of Pretrained Language Model Representations
Yu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang, Jiawei Han
摘要
Topic models have been the prominent tools for automatic topic discovery from text corpora. Despite their effectiveness, topic models suffer from several limitations including the inability of modeling word ordering information in documents, the difficulty of incorporating external linguistic knowledge, and the lack of both accurate and efficient inference methods for approximating the intractable posterior. Recently, pretrained language models (PLMs) have brought astonishing performance improvements to a wide variety of tasks due to their superior representations of text. Interestingly, there have not been standard approaches to deploy PLMs for topic discovery as better alternatives to topic models. In this paper, we begin by analyzing the challenges of using PLM representations for topic discovery, and then propose a joint latent space learning and clustering framework built upon PLM embeddings. In the latent space, topic-word and document-topic distributions are jointly modeled so that the discovered topics can be interpreted by coherent and distinctive terms and meanwhile serve as meaningful summaries of the documents. Our model effectively leverages the strong representation power and superb linguistic features brought by PLMs for topic discovery, and is conceptually simpler than topic models. On two benchmark datasets in different domains, our model generates significantly more coherent and diverse topics than strong topic models, and offers better topic-wise document representations, based on both automatic and human evaluations. 1 CCS CONCEPTS • Information systems → Clustering; Document topic models; • Computing methodologies → Natural language processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- How Do Transformers Learn Topic Structure: Towards a Mechanistic UnderstandingYuchen Li, Yuanzhi Li, Andrej RisteskiICML 2023 · 被引用 87 次
- Heterformer: Transformer-based Deep Node Representation Learning on Heterogeneous Text-Rich NetworksBowen Jin, Yu Zhang, Qi Zhu, Jiawei HanKDD 2023 · 被引用 27 次
- Specious Sites: Tracking the Spread and Sway of Spurious News Stories at ScaleHans W. A. Hanley, Deepak Kumar, Zakir DurumericS&P 2024 · 被引用 18 次
- Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document GenerationYoungjoon Yoo, Jongwon ChoiAAAI 2024 · 被引用 8 次
- Emergence of Hierarchical Emotion Organization in Large Language ModelsMaya Okawa, Bo Zhao, Eric Bigelow, Rose Yu 等ICML 2026 · 被引用 4 次
它引用的顶会 Paper12
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model PretrainingYu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary 等NeurIPS 2021 · 被引用 231 次
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong 等EMNLP 2020 · 被引用 203 次
- Discriminative Topic Mining via Category-Name Guided Text EmbeddingYu Meng, Jiaxin Huang, Guangyuan Wang, Zihan Wang 等WWW 2020 · 被引用 80 次
相关 Paper
- Neural Topic Modeling with Large Language Models in the LoopXiaohao Yang, He Zhao, Weijie Xu, Yuanyuan Qi 等ACL 2025 · 被引用 13 次
- Representing Mixtures of Word Embeddings with Mixtures of Topic EmbeddingsDongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng 等ICLR 2022 · 被引用 56 次
- Explainable and Discourse Topic-aware Neural Language UnderstandingYatin Chaudhary, Hinrich Schütze, Pankaj GuptaICML 2020 · 被引用 7 次
- Beyond prompting: Making Pre-trained Language Models Better Zero-shot Learners by Clustering RepresentationsYu Fei, Zhao Meng, Ping Nie, Roger Wattenhofer 等EMNLP 2022 · 被引用 13 次
- Pre-training and Fine-tuning Neural Topic Model: A Simple yet Effective Approach to Incorporating External KnowledgeLinhai Zhang, Xuemeng Hu, Boyu Wang, Deyu Zhou 等ACL 2022 · 被引用 14 次
