Sparse Parallel Training of Hierarchical Dirichlet Process Topic Models
Alexander Terenin, Måns Magnusson, Leif Jonsson
摘要
To scale non-parametric extensions of probabilistic topic models such as Latent Dirichlet allocation to larger data sets, practitioners rely increasingly on parallel and distributed systems. In this work, we study data-parallel training for the hierarchical Dirichlet process (HDP) topic model. Based upon a representation of certain conditional distributions within an HDP, we propose a doubly sparse data-parallel sampler for the HDP topic model. This sampler utilizes all available sources of sparsity found in natural language-an important way to make computation efficient. We benchmark our method on a well-known corpus (PubMed) with 8m documents and 768m tokens, using a single multi-core machine in under four days.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Tree-Structured Topic Modeling with Nonparametric Neural Variational InferenceZiye Chen, Cheng Ding, Zusheng Zhang, Yanghui Rao 等ACL 2021
- Learning VAE-LDA Models with Rounded Reparameterization TrickRunzhi Tian, Yongyi Mao, Richong ZhangEMNLP 2020 · 被引用 16 次
- Topic Modeling via Full Dependence MixturesDan Fisher, Mark Kozdoba, Shie MannorICML 2020 · 被引用 2 次
- RED-HDP-HMM: Observation-Dependent Durations for Bayesian Nonparametric Sequential ModelsMikołaj Słupiński, Piotr LipinskiICML 2026
- Progressive Tempering Sampler with DiffusionSeveri Rissanen, Ruikang Ouyang, Jiajun He, Wenlin Chen 等ICML 2025
