CEMTM: Contextual Embedding-based Multimodal Topic Modeling
Amirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty, Giuseppe Carenini
Abstract
We introduce CEMTM, a context-enhanced multimodal topic model designed to infer coherent and interpretable topic structures from both short and long documents containing text and images. CEMTM builds on fine-tuned large vision language models (LVLMs) to obtain contextualized embeddings, and employs a distributional attention mechanism to weight token-level contributions to topic inference. A reconstruction objective aligns topic-based representations with the document embedding, encouraging semantic consistency across modalities. Unlike existing approaches, CEMTM can process multiple images per document without repeated encoding and maintains interpretability through explicit word-topic and documenttopic distributions. Extensive experiments on six multimodal benchmarks show that CEMTM consistently outperforms unimodal and multimodal baselines, achieving a remarkable average LLM score of 2.61 (1-3 scale). Further analysis shows its effectiveness in downstream few-shot retrieval and its ability to capture visually grounded semantics in complex domains such as scientific articles 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Sparse Autoencoders are Topic ModelsLeander Girrbach, Zeynep AkataICML 2026 · 2 citations
- Improving Topic Modeling by Distilling Soft Labels from Language ModelsRaymond Li, Amirhossein Abaskohi, Chuyuan Li, Gabriel Murray et al.ICML 2026
Builds on5
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- PromptMTopic: Unsupervised Multimodal Topic Modeling of Memes using Large Language ModelsNirmalendu Prakash, Han Wang, Nguyen-Khoi Hoang, Ming Shan Hee et al.ACM MM 2023 · 24 citations
- A Suite of Generative Tasks for Multi-Level Multimodal Webpage UnderstandingAndrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown et al.EMNLP 2023 · 2 citations
- Neural Attention-Aware Hierarchical Topic ModelYuan Jin, He Zhao, Ming Liu, Lan Du et al.EMNLP 2021
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding TasksZiyan Jiang, Rui Meng, Xinyi Yang, Semih Yavuz et al.ICLR 2025
Related papers
- Seeing to Generalize: How Visual Data Corrects Binding ShortcutsNicolas Buzeta, Felipe del Rio, Cristian Hinostroza, Denis Parra et al.ICML 2026 · 1 citation
- Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document RetrievalDavide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi et al.CVPR 2025
- DREAM: Integrating Hierarchical Multimodal Retrieval with Multi-page Multimodal Language Model for Documents VQAJinxu Zhang, Qiyuan Fan, Yongqi Yu, Yu ZhangACM MM 2025
- A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual CluesYunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding et al.ACL 2023 · 10 citations
- DocLLM: A Layout-Aware Generative Language Model for Multimodal Document UnderstandingDongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma et al.ACL 2024 · 37 citations
