Neural Topic Model via Optimal Transport
He Zhao, Dinh Phung, Viet Huynh, Trung Le, Wray L. Buntine
摘要
Recently, Neural Topic Models (NTMs) inspired by variational autoencoders have obtained increasingly research interest due to their promising results on text analysis. However, it is usually hard for existing NTMs to achieve good document representation and coherent/diverse topics at the same time. Moreover, they often degrade their performance severely on short documents. The requirement of reparameterisation could also comprise their training quality and model flexibility. To address these shortcomings, we present a new neural topic model via the theory of optimal transport (OT). Specifically, we propose to learn the topic distribution of a document by directly minimising its OT distance to the document's word distributions. Importantly, the cost matrix of the OT distance models the weights between topics and words, which is constructed by the distances between topics and words in an embedding space. Our proposed model can be trained efficiently with a differentiable loss. Extensive experiments show that our framework significantly outperforms the state-of-the-art NTMs on discovering more coherent and diverse topics and deriving better document representations for both regular and short texts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic ModelXiaobao Wu, Thong Nguyen, Delvin Zhang, William Yang Wang 等NeurIPS 2024 · 被引用 67 次
- Representing Mixtures of Word Embeddings with Mixtures of Topic EmbeddingsDongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng 等ICLR 2022 · 被引用 56 次
- A Unified Wasserstein Distributional Robustness Framework for Adversarial TrainingAnh Tuan Bui, Trung Le, Quan Hung Tran, He Zhao 等ICLR 2022 · 被引用 54 次
- Tuning Multi-mode Token-level Prompt Alignment across ModalitiesDongsheng Wang, Miaoge Li, Xinyang Liu, Mingsheng Xu 等NeurIPS 2023 · 被引用 49 次
- Transformed Distribution Matching for Missing Value ImputationHe Zhao, Ke Sun, Amir Dezfouli, Edwin V. BonillaICML 2023 · 被引用 46 次
它引用的顶会 Paper1
相关 Paper
- Dynamic Topic Models for Temporal Document NetworksDelvin Ce Zhang, Hady W. LauwICML 2022 · 被引用 26 次
- Neural Topic Modeling with Large Language Models in the LoopXiaohao Yang, He Zhao, Weijie Xu, Yuanyuan Qi 等ACL 2025 · 被引用 13 次
- Neural Attention-Aware Hierarchical Topic ModelYuan Jin, He Zhao, Ming Liu, Lan Du 等EMNLP 2021
- Short Text Topic Modeling with Topic Distribution Quantization and Negative Sampling DecoderXiaobao Wu, Chunping Li, Yan Zhu, Yishu MiaoEMNLP 2020 · 被引用 61 次
- Topic Modeling as Multi-Objective Contrastive OptimizationThong Thanh Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy T. Nguyen 等ICLR 2024 · 被引用 13 次
