Topic Modeling via Full Dependence Mixtures
Dan Fisher, Mark Kozdoba, Shie Mannor
摘要
In this paper we introduce a new approach to topic modelling that scales to large datasets by using a compact representation of the data and by leveraging the GPU architecture. In this approach, topics are learned directly from the co-occurrence data of the corpus. In particular, we introduce a novel mixture model which we term the Full Dependence Mixture (FDM) model. FDMs model second moment under general generative assumptions on the data. While there is previous work on topic modeling using second moments, we develop a direct stochastic optimization procedure for fitting an FDM with a single Kullback Leibler objective. Moment methods in general have the benefit that an iteration no longer needs to scale with the size of the corpus. Our approach allows us to leverage standard optimizers and GPUs for the problem of topic modeling. In particular, we evaluate the approach on two large datasets, NeurIPS papers and a Twitter corpus, with a large number of topics, and show that the approach performs comparably or better than the the standard benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Topic Modeling Revisited: A Document Graph-based Neural Network PerspectiveDazhong Shen, Chuan Qin, Chao Wang, Zheng Dong 等NeurIPS 2021 · 被引用 50 次
- Neural Mixed Counting Models for Dispersed Topic DiscoveryJiemin Wu, Yanghui Rao, Zusheng Zhang, Haoran Xie 等ACL 2020 · 被引用 16 次
- Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN TrainingZuocheng Shi, Jie Sun, Ziyu Song, Mo Sun 等SC 2025 · 被引用 3 次
- Representing Mixtures of Word Embeddings with Mixtures of Topic EmbeddingsDongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng 等ICLR 2022 · 被引用 56 次
- Sparse Parallel Training of Hierarchical Dirichlet Process Topic ModelsAlexander Terenin, Måns Magnusson, Leif JonssonEMNLP 2020
