Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document Generation
Youngjoon Yoo, Jongwon Choi
Abstract
This paper introduces a novel approach for topic modeling utilizing latent codebooks from Vector-Quantized Variational Auto-Encoder (VQ-VAE), discretely encapsulating the rich information of the pre-trained embeddings such as the pre-trained language model. From the novel interpretation of the latent codebooks and embeddings as conceptual bagof-words, we propose a new generative topic model called Topic-VQ-VAE (TVQ-VAE) which inversely generates the original documents related to the respective latent codebook. The TVQ-VAE can visualize the topics with various generative distributions including the traditional BoW distribution and the autoregressive image generation. Our experimental results on document analysis and image generation demonstrate that TVQ-VAE effectively captures the topic context which reveals the underlying structures of the dataset and supports flexible forms of document generation. Official implementation of the proposed TVQ-VAE is available at https: //github.com/clovaai/TVQ-VAE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17f32f11-7e29-44d6-8883-6e79af3e962fCited by top-tier papers2
- Straighten Viscous Rectified Flow via Noise OptimizationJimin Dai, Jiexi Yan, Jian Yang, Lei LuoICCV 2025 · 1 citation
- ArcVQ-VAE: A Spherical Vector Quantization Framework with ArcCosine Additive MarginJaeyung Kim, YoungJoon YooICML 2026
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang et al.ICLR 2022 · 753 citations
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen et al.CVPR 2022 · 607 citations
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image SynthesisPatrick Esser, Robin Rombach, Andreas Blattmann, Björn OmmerNeurIPS 2021 · 187 citations
Related papers
- Neural Attention-Aware Hierarchical Topic ModelYuan Jin, He Zhao, Ming Liu, Lan Du et al.EMNLP 2021
- A Discrete Variational Recurrent Topic Model without the Reparametrization TrickMehdi Rezaee, Francis FerraroNeurIPS 2020 · 31 citations
- SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic QuantizationYuhta Takida, Takashi Shibuya, Wei-Hsiang Liao, Chieh-Hsin Lai et al.ICML 2022 · 99 citations
- Concept-Centric Token Interpretation for Vector-Quantized Generative ModelsTianze Yang, Yucheng Shi, Mengnan Du, Xuansheng Wu et al.ICML 2025
- Diffusion bridges vector quantized variational autoencodersMax Cohen, Guillaume Quispe, Sylvain Le Corff, Charles Ollion et al.ICML 2022 · 16 citations
