Large Language Models Struggle to Describe the Haystack without Human Help: A Social Science-Inspired Evaluation of Topic Models
Zongxia Li, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Paiheng Xu, Daniel Kofi Stephens, Juan Francisco Fung, Alden Dima, Jordan Lee Boyd-Graber
Abstract
A common use of NLP by social scientists is to understand large document collections. Recent data exploration and content analysis have shifted from probabilistic topic models to Large Language Models (LLMs). Yet their effectiveness in helping users understand content in real-world applications remains under explored. This study compares the knowledge users gain from unsupervised LLMs, supervised LLMs, and traditional topic models across two datasets. While unsupervised LLMs generate more human-readable topics, their topics are overly generic for domain-specific datasets and do not help users learn much about the documents. Adding human supervision to LLM generation improves data exploration by mitigating hallucination and over-genericity but requires greater human effort. Traditional topic models, such as Latent Dirichlet Allocation (LDA), remain effective for exploration but are less user-friendly. LLMs struggle to describe the haystack of large corpora without human help, particularly domain-specific data, and face scaling and hallucination limitations due to context length constraints.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7885f66-a006-4157-b564-bca66eec84deBuilds on5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceAlexander Miserlis Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov et al.NeurIPS 2021 · 220 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooMMichelle S. Lam, Janice Teoh, James A. Landay, Jeffrey Heer et al.CHI 2024 · 46 citations
Related papers
- The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?Alexander S. Choi, Syeda Sabrina Akter, JP Singh, Antonios AnastasopoulosEMNLP 2024 · 5 citations
- LLM-Guided Semantic-Aware Clustering for Topic ModelingJianghan Liu, Ziyu Shang, Wenjun Ke, Peng Wang et al.ACL 2025 · 6 citations
- Effects of LLM-based Search on Decision Making: Speed, Accuracy, and OverrelianceSofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, Jake M. HofmanCHI 2025 · 28 citations
- E-LDA: Toward Interpretable LDA Topic Models with Strong Guarantees in Logarithmic Parallel TimeAdam BreuerICML 2025
- LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile AppsShanhui Zhao, Hao Wen, Wenjie Du, Cheng Liang et al.MobiCom 2025 · 6 citations
