M3Hop-CoT: Misogynous Meme Identification with Multimodal Multi-hop Chain-of-Thought
Gitanjali Kumari, Kirtan Jain, Asif Ekbal
Abstract
In recent years, there has been a significant rise in the phenomenon of hate against women on social media platforms, particularly through the use of misogynous memes. These memes often target women with subtle and obscure cues, making their detection a challenging task for automated systems. Recently, Large Language Models (LLMs) have shown promising results in reasoning using Chain-of-Thought (CoT) prompting to generate the intermediate reasoning chains as the rationale to facilitate multimodal tasks, but often neglect cultural diversity and key aspects like emotion and contextual knowledge hidden in the visual modalities. To address this gap, we introduce a Multimodal Multi-hop CoT (M3Hop-CoT) framework for Misogynous meme identification, combining a CLIP-based classifier and a multimodal CoT module with entity-object-relationship integration. M3Hop-CoT employs a three-step multimodal prompting principle to induce emotions, target awareness, and contextual knowledge for meme analysis. Our empirical evaluation, including both qualitative and quantitative analysis, validates the efficacy of the M3Hop-CoT framework on the SemEval-2022 Task 5 (MAMI task) dataset, highlighting its strong performance in the macro-F1 score. Furthermore, we evaluate the model’s generalizability by evaluating it on various benchmark meme datasets, offering a thorough insight into the effectiveness of our approach across different datasets. Codes are available at this link: https://github.com/Gitanjali1801/LLM_CoT
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4072703e-b19e-41e5-884e-2194199a8e2fCited by top-tier papers5
- MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion RecognitionJian Chen, Yuxuan Hu, Haifeng Lu, Wei Wang et al.ACM MM 2025 · 5 citations
- Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionFengjun Pan, Xiaobao Wu, Tho Quan, Anh Tuan LuuWWW 2026 · 2 citations
- MemeIntel: Explainable Detection of Propagandistic and Hateful MemesMohamed Bayan Kmainasi, Abul Hasnat, Md. Arid Hasan, Ali Ezzat Shahroor et al.EMNLP 2025 · 1 citation
- MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language ModelsZixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo et al.EMNLP 2025
- AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on HarmfulnessZixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo et al.ACL 2025
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
Related papers
- BANMIME : Misogyny Detection with Metaphor Explanation on Bangla MemesMd Ayon Mia, Akm Moshiur Rahman Mazumder, Khadiza Sultana Sayma, Md Fahim et al.EMNLP 2025
- Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme DetectionJingbiao Mei, Jinghong Chen, Guangyu Yang, Weizhe Lin et al.EMNLP 2025 · 2 citations
- On the Evolution of (Hateful) Memes by Means of Multimodal Contrastive LearningYiting Qu, Xinlei He, Shannon Pierson, Michael Backes et al.S&P 2023
- I know what you MEME! Understanding and Detecting Harmful Memes with Multimodal Large Language ModelsYong Zhuang, Keyan Guo, Juan Wang, Yiheng Jing et al.NDSS 2025
- Ask, Acquire, Understand: A Multimodal Agent-based Framework for Social Abuse Detection in MemesXuanrui Lin, Chao Jia, Junhui Ji, Hui Han et al.WWW 2025 · 9 citations
