Representing Mixtures of Word Embeddings with Mixtures of Topic Embeddings
Dongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng, Korawat Tanwisuth, Bo Chen, Mingyuan Zhou
摘要
A topic model is often formulated as a generative model that explains how each word of a document is generated given a set of topics and document-specific topic proportions. It is focused on capturing the word co-occurrences in a document and hence often suffers from poor performance in analyzing short documents. In addition, its parameter estimation often relies on approximate posterior inference that is either not scalable or suffering from large approximation error. This paper introduces a new topic-modeling framework where each document is viewed as a set of word embedding vectors and each topic is modeled as an embedding vector in the same embedding space. Embedding the words and topics in the same vector space, we define a method to measure the semantic difference between the embedding vectors of the words of a document and these of the topics, and optimize the topic embeddings to minimize the expected difference over all documents. Experiments on text analysis demonstrate that the proposed method, which is amenable to mini-batch stochastic gradient descent based optimization and hence scalable to big corpora, provides competitive performance in discovering more coherent and diverse topics and extracting better document representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Effective Neural Topic Modeling with Embedding Clustering RegularizationXiaobao Wu, Xinshuai Dong, Thong Thanh Nguyen, Anh Tuan LuuICML 2023 · 被引用 87 次
- FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic ModelXiaobao Wu, Thong Nguyen, Delvin Zhang, William Yang Wang 等NeurIPS 2024 · 被引用 67 次
- Enhancing Minority Classes by Mixing: An Adaptative Optimal Transport Approach for Long-tailed ClassificationJintong Gao, He Zhao, Zhuo Li, Dandan GuoNeurIPS 2023 · 被引用 64 次
- Tuning Multi-mode Token-level Prompt Alignment across ModalitiesDongsheng Wang, Miaoge Li, Xinyang Liu, Mingsheng Xu 等NeurIPS 2023 · 被引用 49 次
- Transformed Distribution Matching for Missing Value ImputationHe Zhao, Ke Sun, Amir Dezfouli, Edwin V. BonillaICML 2023 · 被引用 46 次
它引用的顶会 Paper7
- A Prototype-Oriented Framework for Unsupervised Domain AdaptationKorawat Tanwisuth, Xinjie Fan, Huangjie Zheng, Shujian Zhang 等NeurIPS 2021 · 被引用 136 次
- Neural Topic Model via Optimal TransportHe Zhao, Dinh Phung, Viet Huynh, Trung Le 等ICLR 2021 · 被引用 100 次
- Sawtooth Factorial Topic Embeddings Guided Gamma Belief NetworkZhibin Duan, Dongsheng Wang, Bo Chen, Chaojie Wang 等ICML 2021 · 被引用 49 次
- OTLDA: A Geometry-aware Optimal Transport Approach for Topic ModelingViet Huynh, He Zhao, Dinh PhungNeurIPS 2020 · 被引用 29 次
- Recurrent Hierarchical Topic-Guided RNN for Language GenerationDandan Guo, Bo Chen, Ruiying Lu, Mingyuan ZhouICML 2020 · 被引用 20 次
相关 Paper
- Topic Discovery via Latent Space Clustering of Pretrained Language Model RepresentationsYu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang 等WWW 2022 · 被引用 73 次
- Visualizing Temporal Topic Embeddings with a CompassDaniel Palamarchuk, Lemara Williams, Brian Mayer, Thomas Danielson 等IEEE VIS 2024 · 被引用 3 次
- Meta-Complementing the Semantics of Short Texts in Neural Topic ModelsDelvin Ce Zhang, Hady W. LauwNeurIPS 2022 · 被引用 10 次
- TAN-NTM: Topic Attention Networks for Neural Topic ModelingMadhur Panwar, Shashank Shailabh, Milan Aggarwal, Balaji KrishnamurthyACL 2021
- Adaptive Pseudo-Labeling via Word Coherence for Topic ModelingBohan Yoon, Hyejin JangKDD 2026
