Is Automated Topic Model Evaluation Broken? The Incoherence of Coherence
Alexander Miserlis Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov, Jordan L. Boyd-Graber, Philip Resnik
Abstract
Topic model evaluation, like evaluation of other unsupervised methods, can be contentious. However, the field has coalesced around automated estimates of topic coherence, which rely on the frequency of word co-occurrences in a reference corpus. Contemporary neural topic models surpass classical ones according to these metrics. At the same time, topic model evaluation suffers from a validation gap: automated coherence, developed for classical models, has not been validated using human experimentation for neural models. In addition, a meta-analysis of topic modeling literature reveals a substantial standardization gap in automated topic modeling benchmarks. To address the validation gap, we compare automated coherence with the two most widely accepted human judgment tasks: topic rating and word intrusion. To address the standardization gap, we systematically evaluate a dominant classical model and two state-of-the-art neural models on two commonly used datasets. Automated evaluations declare a winning model when corresponding human evaluations do not, calling into question the validity of fully automatic evaluations independent of human judgments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 919df600-6b0f-40f4-b6c8-638cabf09340Cited by top-tier papers18
- FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic ModelXiaobao Wu, Thong Nguyen, Delvin Zhang, William Yang Wang et al.NeurIPS 2024 · 67 citations
- PromptMTopic: Unsupervised Multimodal Topic Modeling of Memes using Large Language ModelsNirmalendu Prakash, Han Wang, Nguyen-Khoi Hoang, Ming Shan Hee et al.ACM MM 2023 · 24 citations
- Large-Scale Correlation Analysis of Automated Metrics for Topic ModelsJia Peng Lim, Hady W. LauwACL 2023 · 13 citations
- Topic Modeling as Multi-Objective Contrastive OptimizationThong Thanh Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy T. Nguyen et al.ICLR 2024 · 13 citations
- Topic Modeling With Topological Data AnalysisCiarán Byrne, Danijela Horak, Karo Moilanen, Amandla MabonaEMNLP 2022 · 7 citations
Builds on9
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia et al.EMNLP 2020 · 76 citations
- Short Text Topic Modeling with Topic Distribution Quantization and Negative Sampling DecoderXiaobao Wu, Chunping Li, Yan Zhu, Yishu MiaoEMNLP 2020 · 61 citations
- Graph Attention Topic Modeling NetworkLiang Yang, Fan Wu, Junhua Gu, Chuan Wang et al.WWW 2020 · 57 citations
- A Discrete Variational Recurrent Topic Model without the Reparametrization TrickMehdi Rezaee, Francis FerraroNeurIPS 2020 · 31 citations
- Neural Topic Modeling with Cycle-Consistent Adversarial TrainingXuemeng Hu, Rui Wang, Deyu Zhou, Yuxuan XiongEMNLP 2020 · 25 citations
Related papers
- Evaluation of Thematic Coherence in MicroblogsIman Munire Bilal, Bo Wang, Maria Liakata, Rob Procter et al.ACL 2021
- Effective Neural Topic Modeling with Embedding Clustering RegularizationXiaobao Wu, Xinshuai Dong, Thong Thanh Nguyen, Anh Tuan LuuICML 2023 · 87 citations
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu et al.EMNLP 2020 · 3 citations
- Uncovering Competency Gaps in Large Language Models and Their BenchmarksMaty Bohacek, Nino Scherrer, Nicholas Dufour, Thomas Leung et al.ICML 2026
- Neural Topic Modeling with Bidirectional Adversarial TrainingRui Wang, Xuemeng Hu, Deyu Zhou, Yulan He et al.ACL 2020 · 77 citations
