PromptMTopic: Unsupervised Multimodal Topic Modeling of Memes using Large Language Models
Nirmalendu Prakash, Han Wang, Nguyen-Khoi Hoang, Ming Shan Hee, Roy Ka-Wei Lee
Abstract
The proliferation of social media has given rise to a new form of communication: memes. Memes are multimodal and often contain a combination of text and visual elements that convey meaning, humor, and cultural significance. While meme analysis has been an active area of research, little work has been done on unsupervised multimodal topic modeling of memes, which is important for content moderation, social media analysis, and cultural studies. We propose PromptMTopic, a novel multimodal prompt-based model designed to learn topics from both text and visual modalities by leveraging the language modeling capabilities of large language models. Our model effectively extracts and clusters topics learned from memes, considering the semantic interaction between the text and visual modalities. We evaluate our proposed model through extensive experiments on three real-world meme datasets, which demonstrate its superiority over state-of-the-art topic modeling baselines in learning descriptive topics in memes. Additionally, our qualitative analysis shows that PromptMTopic can identify meaningful and culturally relevant topics from memes. Our work contributes to the understanding of the topics and themes of memes, a crucial form of communication in today's society. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.
• Computing methodologies → Natural language processing; Computer vision representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37b1f9a2-6ab2-41e6-b8d8-9eea239023abCited by top-tier papers3
- ArMeme: Propagandistic Content in Arabic MemesFiroj Alam, Abul Hasnat, Fatema Ahmad, Md. Arid Hasan et al.EMNLP 2024 · 4 citations
- Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal TriggersRuofei Wang, Hongzhan Lin, Ziyuan Luo, Ka Chun Cheung et al.AAAI 2025
- CEMTM: Contextual Embedding-based Multimodal Topic ModelingAmirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty et al.EMNLP 2025
Builds on6
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceAlexander Miserlis Hoyle, Pranav Goel, Andrew Hian-Cheong, Denis Peskov et al.NeurIPS 2021 · 220 citations
- Disentangling Hate in Online MemesRoy Ka-Wei Lee, Rui Cao, Ziqing Fan, Jing Jiang et al.ACM MM 2021 · 85 citations
- Prompting for Multimodal Hateful Meme ClassificationRui Cao, Roy Ka-Wei Lee, Wen-Haw Chong, Jing JiangEMNLP 2022 · 68 citations
Related papers
- Ask, Acquire, Understand: A Multimodal Agent-based Framework for Social Abuse Detection in MemesXuanrui Lin, Chao Jia, Junhui Ji, Hui Han et al.WWW 2025 · 9 citations
- Pro-Cap: Leveraging a Frozen Vision-Language Model for Hateful Meme DetectionRui Cao, Ming Shan Hee, Adriel Kuek, Wen-Haw Chong et al.ACM MM 2023 · 54 citations
- Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal FusionYi Shi, Wenlong Meng, Zhenyuan Guo, Chengkun Wei et al.AAAI 2026
- Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionFengjun Pan, Xiaobao Wu, Tho Quan, Anh Tuan LuuWWW 2026 · 2 citations
- MetaGPT: A Large Vision-Language Model for Meme Metaphor UnderstandingBo Xu, Chenyuan Wang, Xinyu Chen, Hongfei Lin et al.AAAI 2026
