Pre-training and Fine-tuning Neural Topic Model: A Simple yet Effective Approach to Incorporating External Knowledge
Linhai Zhang, Xuemeng Hu, Boyu Wang, Deyu Zhou, Qian-Wen Zhang, Yunbo Cao
摘要
Recent years have witnessed growing interests in incorporating external knowledge such as pre-trained word embeddings (PWEs) or pre-trained language models (PLMs) into neural topic modeling. However, we found that employing PWEs and PLMs for topic modeling only achieved limited performance improvements but with huge computational overhead. In this paper, we propose a novel strategy to incorporate external knowledge into neural topic modeling where the neural topic model is pre-trained on a large corpus and then fine-tuned on the target dataset. Experiments have been conducted on three datasets and results show that the proposed approach significantly outperforms both current state-of-the-art neural topic models and some topic modeling approaches enhanced with PWEs or PLMs. Moreover, further study shows that the proposed approach greatly reduces the need for the huge size of training data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Topic Modeling as Multi-Objective Contrastive OptimizationThong Thanh Nguyen, Xiaobao Wu, Xinshuai Dong, Cong-Duy T. Nguyen 等ICLR 2024 · 被引用 13 次
- Disentangling Transformer Language Models as Superposed Topic ModelsJia Peng Lim, Hady W. LauwEMNLP 2023 · 被引用 2 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Neural Topic Modeling with Cycle-Consistent Adversarial TrainingXuemeng Hu, Rui Wang, Deyu Zhou, Yuxuan XiongEMNLP 2020 · 被引用 25 次
- CluBERT: A Cluster-Based Approach for Learning Sense Distributions in Multiple LanguagesTommaso Pasini, Federico Scozzafava, Bianca ScarliniACL 2020 · 被引用 23 次
- Improving Neural Topic Models using Knowledge DistillationAlexander Miserlis Hoyle, Pranav Goel, Philip ResnikEMNLP 2020 · 被引用 5 次
相关 Paper
- Neural Attention-Aware Hierarchical Topic ModelYuan Jin, He Zhao, Ming Liu, Lan Du 等EMNLP 2021
- Neural Topic Modeling with Large Language Models in the LoopXiaohao Yang, He Zhao, Weijie Xu, Yuanyuan Qi 等ACL 2025 · 被引用 13 次
- Topic Discovery via Latent Space Clustering of Pretrained Language Model RepresentationsYu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang 等WWW 2022 · 被引用 73 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 被引用 4 次
