FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic Model
Xiaobao Wu, Thong Nguyen, Delvin Zhang, William Yang Wang, Anh Tuan Luu
摘要
Topic models have been evolving rapidly over the years, from conventional to recent neural models. However, existing topic models generally struggle with either effectiveness, efficiency, or stability, highly impeding their practical applications. In this paper, we propose FASTopic, a fast, adaptive, stable, and transferable topic model. FASTopic follows a new paradigm: Dual Semantic-relation Reconstruction (DSR). Instead of previous conventional, VAE-based, or clustering-based methods, DSR directly models the semantic relations among document embeddings from a pretrained Transformer and learnable topic and word embeddings. By reconstructing through these semantic relations, DSR discovers latent topics. This brings about a neat and efficient topic modeling framework. We further propose a novel Embedding Transport Plan (ETP) method. Rather than early straightforward approaches, ETP explicitly regularizes the semantic relations as optimal transport plans. This addresses the relation bias issue and thus leads to effective topic modeling. Extensive experiments on benchmark datasets demonstrate that our FASTopic shows superior effectiveness, efficiency, adaptivity, stability, and transferability, compared to state-of-the-art baselines across various scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World KnowledgeXiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou 等ACL 2025 · 被引用 35 次
- LLM-Guided Semantic-Aware Clustering for Topic ModelingJianghan Liu, Ziyu Shang, Wenjun Ke, Peng Wang 等ACL 2025 · 被引用 6 次
- Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionFengjun Pan, Xiaobao Wu, Tho Quan, Anh Tuan LuuWWW 2026 · 被引用 2 次
- Constrained Non-negative Matrix Factorization for Guided Topic Modeling of Minority TopicsSeyedeh Fatemeh Ebrahimi, Jaakko PeltonenEMNLP 2025 · 被引用 2 次
- Sparse Autoencoders are Topic ModelsLeander Girrbach, Zeynep AkataICML 2026 · 被引用 2 次
它引用的顶会 Paper18
- Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceAlexander Miserlis Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov 等NeurIPS 2021 · 被引用 220 次
- Neural Topic Model via Optimal TransportHe Zhao, Dinh Phung, Viet Huynh, Trung Le 等ICLR 2021 · 被引用 100 次
- Effective Neural Topic Modeling with Embedding Clustering RegularizationXiaobao Wu, Xinshuai Dong, Thong Thanh Nguyen, Anh Tuan LuuICML 2023 · 被引用 87 次
- Short Text Topic Modeling with Topic Distribution Quantization and Negative Sampling DecoderXiaobao Wu, Chunping Li, Yan Zhu, Yishu MiaoEMNLP 2020 · 被引用 61 次
- Representing Mixtures of Word Embeddings with Mixtures of Topic EmbeddingsDongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng 等ICLR 2022 · 被引用 56 次
相关 Paper
- Dynamic Topic Models for Temporal Document NetworksDelvin Ce Zhang, Hady W. LauwICML 2022 · 被引用 26 次
- Contrastive Learning for Neural Topic ModelThong Nguyen, Anh Tuan LuuNeurIPS 2021 · 被引用 82 次
- Neural Attention-Aware Hierarchical Topic ModelYuan Jin, He Zhao, Ming Liu, Lan Du 等EMNLP 2021
- Adaptive Pseudo-Labeling via Word Coherence for Topic ModelingBohan Yoon, Hyejin JangKDD 2026
- OTLDA: A Geometry-aware Optimal Transport Approach for Topic ModelingViet Huynh, He Zhao, Dinh PhungNeurIPS 2020 · 被引用 29 次
