Sparse Parallel Training of Hierarchical Dirichlet Process Topic Models
Alexander Terenin, Måns Magnusson, Leif Jonsson
Abstract
To scale non-parametric extensions of probabilistic topic models such as Latent Dirichlet allocation to larger data sets, practitioners rely increasingly on parallel and distributed systems. In this work, we study data-parallel training for the hierarchical Dirichlet process (HDP) topic model. Based upon a representation of certain conditional distributions within an HDP, we propose a doubly sparse data-parallel sampler for the HDP topic model. This sampler utilizes all available sources of sparsity found in natural language-an important way to make computation efficient. We benchmark our method on a well-known corpus (PubMed) with 8m documents and 768m tokens, using a single multi-core machine in under four days.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Tree-Structured Topic Modeling with Nonparametric Neural Variational InferenceZiye Chen, Cheng Ding, Zusheng Zhang, Yanghui Rao et al.ACL 2021
- Learning VAE-LDA Models with Rounded Reparameterization TrickRunzhi Tian, Yongyi Mao, Richong ZhangEMNLP 2020 · 16 citations
- Topic Modeling via Full Dependence MixturesDan Fisher, Mark Kozdoba, Shie MannorICML 2020 · 2 citations
- RED-HDP-HMM: Observation-Dependent Durations for Bayesian Nonparametric Sequential ModelsMikołaj Słupiński, Piotr LipinskiICML 2026
- Progressive Tempering Sampler with DiffusionSeveri Rissanen, Ruikang Ouyang, Jiajun He, Wenlin Chen et al.ICML 2025
