Neural Attention-Aware Hierarchical Topic Model
Yuan Jin, He Zhao, Ming Liu, Lan Du, Wray L. Buntine
Abstract
Neural topic models (NTMs) apply deep neural networks to topic modelling. Despite their success, NTMs generally ignore two important aspects: (1) only document-level word count information is utilized for the training, while more fine-grained sentence-level information is ignored, and (2) external semantic knowledge regarding documents, sentences and words are not exploited for the training. To address these issues, we propose a variational autoencoder (VAE) NTM model that jointly reconstructs the sentence and document word counts using combinations of bag-of-words (BoW) topical embeddings and pre-trained semantic embeddings. The pre-trained embeddings are first transformed into a common latent topical space to align their semantics with the BoW embeddings. Our model also features hierarchical KL divergence to leverage embeddings of each document to regularize those of their sentences, thereby paying more attention to semantically relevant sentences. Both quantitative and qualitative experiments have shown the efficacy of our model in 1) lowering the reconstruction errors at both the sentence and document levels, and 2) discovering more coherent topics from real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- ConvNTM: Conversational Neural Topic ModelHongda Sun, Quan Tu, Jinpeng Li, Rui YanAAAI 2023 · 6 citations
- CEMTM: Contextual Embedding-based Multimodal Topic ModelingAmirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty et al.EMNLP 2025
Builds on2
Related papers
- Do sequence-to-sequence VAEs learn global features of sentences?Tom Bosc, Pascal VincentEMNLP 2020 · 5 citations
- TAN-NTM: Topic Attention Networks for Neural Topic ModelingMadhur Panwar, Shashank Shailabh, Milan Aggarwal, Balaji KrishnamurthyACL 2021
- Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document GenerationYoungjoon Yoo, Jongwon ChoiAAAI 2024 · 8 citations
- Neural Topic Model via Optimal TransportHe Zhao, Dinh Phung, Viet Huynh, Trung Le et al.ICLR 2021 · 100 citations
- TopicNet: Semantic Graph-Guided Topic DiscoveryZhibin Duan, Yishi Xu, Bo Chen, Dongsheng Wang et al.NeurIPS 2021 · 19 citations
