E-LDA: Toward Interpretable LDA Topic Models with Strong Guarantees in Logarithmic Parallel Time
Adam Breuer
摘要
In this paper, we provide the first practical algorithms with provable guarantees for the problem of inferring the topics assigned to each document in an LDA topic model. This is the primary inference problem for many applications of topic models in social science, data exploration, and causal inference settings. We obtain this result by showing a novel non-gradient-based, combinatorial approach to estimating topic models. This yields algorithms that converge to near-optimal posterior probability in logarithmic parallel computation time (adaptivity)-exponentially faster than any known LDA algorithm. We also show that our approach can provide interpretability guarantees such that each learned topic is formally associated with a known keyword. Finally, we show that unlike alternatives, our approach can maintain the independence assumptions necessary to use the learned topic model for downstream causal inference methods that allow researchers to study topics as treatments. In terms of practical performance, our approach consistently returns solutions of higher semantic quality than solutions from state-of-the-art LDA algorithms, neural topic models, and LLMbased topic models across a diverse range of text datasets and evaluation parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Effective Neural Topic Modeling with Embedding Clustering RegularizationXiaobao Wu, Xinshuai Dong, Thong Thanh Nguyen, Anh Tuan LuuICML 2023 · 被引用 87 次
- FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic ModelXiaobao Wu, Thong Nguyen, Delvin Zhang, William Yang Wang 等NeurIPS 2024 · 被引用 67 次
- The FAST Algorithm for Submodular MaximizationAdam Breuer, Eric Balkanski, Yaron SingerICML 2020 · 被引用 39 次
- On the Unreasonable Effectiveness of the Greedy Algorithm: Greedy Adapts to SharpnessSebastian Pokutta, Mohit Singh, Alfredo TorricoICML 2020 · 被引用 12 次
相关 Paper
- Topic Modeling Revisited: A Document Graph-based Neural Network PerspectiveDazhong Shen, Chuan Qin, Chao Wang, Zheng Dong 等NeurIPS 2021 · 被引用 50 次
- Large Language Models Struggle to Describe the Haystack without Human Help: A Social Science-Inspired Evaluation of Topic ModelsZongxia Li, Lorena Calvo-Bartolomé, Alexander Miserlis Hoyle, Paiheng Xu 等ACL 2025
- Beyond Labels and Topics: Discovering Causal Relationships in Neural Topic ModelingYi-Kun Tang, Heyan Huang, Xuewen Shi, Xian-Ling MaoWWW 2024 · 被引用 4 次
- Neural Topic Modeling with Bidirectional Adversarial TrainingRui Wang, Xuemeng Hu, Deyu Zhou, Yulan He 等ACL 2020 · 被引用 77 次
- EvaLDA: Efficient Evasion Attacks Towards Latent Dirichlet AllocationQi Zhou, Haipeng Chen, Yitao Zheng, Zhen WangAAAI 2021 · 被引用 5 次
