CEMTM: Contextual Embedding-based Multimodal Topic Modeling
Amirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty, Giuseppe Carenini
摘要
We introduce CEMTM, a context-enhanced multimodal topic model designed to infer coherent and interpretable topic structures from both short and long documents containing text and images. CEMTM builds on fine-tuned large vision language models (LVLMs) to obtain contextualized embeddings, and employs a distributional attention mechanism to weight token-level contributions to topic inference. A reconstruction objective aligns topic-based representations with the document embedding, encouraging semantic consistency across modalities. Unlike existing approaches, CEMTM can process multiple images per document without repeated encoding and maintains interpretability through explicit word-topic and documenttopic distributions. Extensive experiments on six multimodal benchmarks show that CEMTM consistently outperforms unimodal and multimodal baselines, achieving a remarkable average LLM score of 2.61 (1-3 scale). Further analysis shows its effectiveness in downstream few-shot retrieval and its ability to capture visually grounded semantics in complex domains such as scientific articles 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Sparse Autoencoders are Topic ModelsLeander Girrbach, Zeynep AkataICML 2026 · 被引用 2 次
- Improving Topic Modeling by Distilling Soft Labels from Language ModelsRaymond Li, Amirhossein Abaskohi, Chuyuan Li, Gabriel Murray 等ICML 2026
它引用的顶会 Paper5
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- PromptMTopic: Unsupervised Multimodal Topic Modeling of Memes using Large Language ModelsNirmalendu Prakash, Han Wang, Nguyen-Khoi Hoang, Ming Shan Hee 等ACM MM 2023 · 被引用 24 次
- A Suite of Generative Tasks for Multi-Level Multimodal Webpage UnderstandingAndrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown 等EMNLP 2023 · 被引用 2 次
- Neural Attention-Aware Hierarchical Topic ModelYuan Jin, He Zhao, Ming Liu, Lan Du 等EMNLP 2021
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding TasksZiyan Jiang, Rui Meng, Xinyi Yang, Semih Yavuz 等ICLR 2025
相关 Paper
- Seeing to Generalize: How Visual Data Corrects Binding ShortcutsNicolas Buzeta, Felipe del Rio, Cristian Hinostroza, Denis Parra 等ICML 2026 · 被引用 1 次
- Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document RetrievalDavide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 等CVPR 2025
- DREAM: Integrating Hierarchical Multimodal Retrieval with Multi-page Multimodal Language Model for Documents VQAJinxu Zhang, Qiyuan Fan, Yongqi Yu, Yu ZhangACM MM 2025
- A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual CluesYunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding 等ACL 2023 · 被引用 10 次
- DocLLM: A Layout-Aware Generative Language Model for Multimodal Document UnderstandingDongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma 等ACL 2024 · 被引用 37 次
