Neural Topic Model via Optimal Transport
He Zhao, Dinh Phung, Viet Huynh, Trung Le, Wray L. Buntine
Abstract
Recently, Neural Topic Models (NTMs) inspired by variational autoencoders have obtained increasingly research interest due to their promising results on text analysis. However, it is usually hard for existing NTMs to achieve good document representation and coherent/diverse topics at the same time. Moreover, they often degrade their performance severely on short documents. The requirement of reparameterisation could also comprise their training quality and model flexibility. To address these shortcomings, we present a new neural topic model via the theory of optimal transport (OT). Specifically, we propose to learn the topic distribution of a document by directly minimising its OT distance to the document's word distributions. Importantly, the cost matrix of the OT distance models the weights between topics and words, which is constructed by the distances between topics and words in an embedding space. Our proposed model can be trained efficiently with a differentiable loss. Extensive experiments show that our framework significantly outperforms the state-of-the-art NTMs on discovering more coherent and diverse topics and deriving better document representations for both regular and short texts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e148596f-aa2f-41c9-a2a6-efb3ac895d39Cited by top-tier papers26
- FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic ModelXiaobao Wu, Thong Nguyen, Delvin Zhang, William Yang Wang et al.NeurIPS 2024 · 67 citations
- Representing Mixtures of Word Embeddings with Mixtures of Topic EmbeddingsDongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng et al.ICLR 2022 · 56 citations
- A Unified Wasserstein Distributional Robustness Framework for Adversarial TrainingAnh Tuan Bui, Trung Le, Quan Hung Tran, He Zhao et al.ICLR 2022 · 54 citations
- Tuning Multi-mode Token-level Prompt Alignment across ModalitiesDongsheng Wang, Miaoge Li, Xinyang Liu, Mingsheng Xu et al.NeurIPS 2023 · 49 citations
- Transformed Distribution Matching for Missing Value ImputationHe Zhao, Ke Sun, Amir Dezfouli, Edwin V. BonillaICML 2023 · 46 citations
Builds on1
Related papers
- Dynamic Topic Models for Temporal Document NetworksDelvin Ce Zhang, Hady W. LauwICML 2022 · 26 citations
- Neural Topic Modeling with Large Language Models in the LoopXiaohao Yang, He Zhao, Weijie Xu, Yuanyuan Qi et al.ACL 2025 · 13 citations
- Neural Attention-Aware Hierarchical Topic ModelYuan Jin, He Zhao, Ming Liu, Lan Du et al.EMNLP 2021
- Short Text Topic Modeling with Topic Distribution Quantization and Negative Sampling DecoderXiaobao Wu, Chunping Li, Yan Zhu, Yishu MiaoEMNLP 2020 · 61 citations
- Topic Modeling as Multi-Objective Contrastive OptimizationThong Thanh Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy T. Nguyen et al.ICLR 2024 · 13 citations
