Large Language Models Struggle to Describe the Haystack without Human Help: A Social Science-Inspired Evaluation of Topic Models
Zongxia Li, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Paiheng Xu, Daniel Kofi Stephens, Juan Francisco Fung, Alden Dima, Jordan Lee Boyd-Graber
摘要
A common use of NLP by social scientists is to understand large document collections. Recent data exploration and content analysis have shifted from probabilistic topic models to Large Language Models (LLMs). Yet their effectiveness in helping users understand content in real-world applications remains under explored. This study compares the knowledge users gain from unsupervised LLMs, supervised LLMs, and traditional topic models across two datasets. While unsupervised LLMs generate more human-readable topics, their topics are overly generic for domain-specific datasets and do not help users learn much about the documents. Adding human supervision to LLM generation improves data exploration by mitigating hallucination and over-genericity but requires greater human effort. Traditional topic models, such as Latent Dirichlet Allocation (LDA), remain effective for exploration but are less user-friendly. LLMs struggle to describe the haystack of large corpora without human help, particularly domain-specific data, and face scaling and hallucination limitations due to context length constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceAlexander Miserlis Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov 等NeurIPS 2021 · 被引用 220 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooMMichelle S. Lam, Janice Teoh, James A. Landay, Jeffrey Heer 等CHI 2024 · 被引用 46 次
相关 Paper
- The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?Alexander S. Choi, Syeda Sabrina Akter, JP Singh, Antonios AnastasopoulosEMNLP 2024 · 被引用 5 次
- LLM-Guided Semantic-Aware Clustering for Topic ModelingJianghan Liu, Ziyu Shang, Wenjun Ke, Peng Wang 等ACL 2025 · 被引用 6 次
- Effects of LLM-based Search on Decision Making: Speed, Accuracy, and OverrelianceSofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, Jake M. HofmanCHI 2025 · 被引用 28 次
- E-LDA: Toward Interpretable LDA Topic Models with Strong Guarantees in Logarithmic Parallel TimeAdam BreuerICML 2025
- LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile AppsShanhui Zhao, Hao Wen, Wenjie Du, Cheng Liang 等MobiCom 2025 · 被引用 6 次
