Exploring Chain-of-Thought for Multi-modal Metaphor Detection
Yanzhi Xu, Yueying Hua, Shichen Li, Zhongqing Wang
Abstract
Metaphors are commonly found in advertising and internet memes. However, the free form of internet memes often leads to a lack of high-quality textual data. Metaphor identification demands a deep interpretation of both textual and visual elements, requiring extensive common-sense knowledge, which poses a challenge to language models. To address these challenges, we propose a compact framework that enhances the small model by distilling knowledge from Multi-modal Large Language Models(MLLMS). Specifically, our approach designs a three-step process inspired by Chain-of-Thought (CoT) that extracts and integrates knowledge from larger models into smaller ones. We also developed a modality fusion architecture to transform knowledge from large models into metaphor features, supplemented by auxiliary tasks to improve model performance. Experimental results on the MET-MEME dataset demonstrate that our method not only effectively enhances the metaphor identification capabilities of small models but also outperforms existing models. To our knowledge, this is the first systematic study leveraging MLLMs in metaphor identification tasks. modality identification, multi-modal metaphor 044 identification not only spots metaphors in sentences 045 but also categorizes them as image-dominated, text-046 dominated, or complementary. The second major 047 challenge arises from the poor quality of textual 048 content, mainly sourced from advertisements and 049 memes on social media. Texts give the image more 050 metaphorical features. Recent efforts use OCR (Op-051 tical Character identification) to extract texts in the 052 image. However, only relying on OCR to convert 053 them into parallel texts leads to the loss of texts' 054 positional information. Figure 1 presents a repre-055 sentative example, symbolizing how 'PUBG' (a 056 video game) acts like a trap preventing 'me' from 057 achieving my 'life goals'. 058 To overcome these challenges, we hope to gain 059 126 Zhang et al., 2023a). Unlike the aforementioned 127 approaches that extract information from different 128 modalities and directly merge them, we leverage 129 LLMs employing the CoT method to analyze fea-130 tures between modalities, aiding downstream mod-131 els in cross-modal fusion. 132 3 Method 133 We propose a novel framework based on knowledge 134 distillation from MLLMs to enhance metaphor 135 identification. In this section we first introduce 136 the task definition(3.1) and the complete model 137 architecture((3.2). After that, we elaborate on 138 knowledge acquisition from MLLMs using the CoT 139 method(3.3) and the implementation of the down-140 stream fusion module(3.4). Finally, we provide a 141 brief exposition of the training methodology (3.5). 142 3.1 Task Definition 143 Formally, the task of multi-modal metaphor iden-144 tification falls under the typical category of multi-145 modal classification problems. Given a set of cross-146 modal sample pairs, the task aims to determine 147 whether metaphorical features are present and pro-148 vide a classification result. Our work focuses on 149 the identification of metaphors in image-text pairs, 150 thus the task is represented as: 151 Y = F (x I , x T ) (1) 152 where x I and x T respectively denote the features 153 of the image and text modalities. Our objective is 154 to utilize a more effective method F to ensure that 155 the classification result Ŷ more closely aligns with 156 the true value y. 157 372 due to their strong performance in both Chinese 373 and English corpora. We fine-tuned both models 374 separately using LoRA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dd4008e-c0ea-41d3-9d02-56921ebdf902Cited by top-tier papers12
- AdaReasoner: Adaptive Reasoning Enables More Flexible ThinkingXiangqi Wang, Yue Huang, Yanbo Wang, Xiaonan Luo et al.NeurIPS 2025 · 21 citations
- Cultural Bias Matters: A Cross-Cultural Benchmark Dataset and Sentiment-Enriched Model for Understanding Multimodal MetaphorsSenqi Yang, Dongyu Zhang, Jing Ren, Ziqi Xu et al.ACL 2025 · 11 citations
- Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor ProcessingFengying Ye, Shanshan Wang, Lidia S. Chao, Derek F. WongACL 2026 · 7 citations
- Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion KnowledgeLi Zheng, Sihang Wang, Hao Fei, Zuquan Peng et al.ACL 2025 · 5 citations
- AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language ReasoningXiping Li, Jianghong MaACL 2026 · 3 citations
Builds on4
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
- Modeling Conceptual Attribute Likeness and Domain Inconsistency for Metaphor DetectionYuan Tian, Nan Xu, Wenji Mao, Daniel ZengEMNLP 2023 · 4 citations
- Verb Metaphor Detection via Contextual Relation LearningWei Song, Shuhui Zhou, Ruiji Fu, Ting Liu et al.ACL 2021
Related papers
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He et al.ICCV 2025 · 9 citations
- IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-IdentificationYuhao Wang, Yongfeng Lv, Pingping Zhang, Huchuan LuCVPR 2025
- Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal FusionYi Shi, Wenlong Meng, Zhenyuan Guo, Chengkun Wei et al.AAAI 2026
- Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme DetectionFengjun Pan, Xiaobao Wu, Tho Quan, Anh Tuan LuuWWW 2026 · 2 citations
- Miko: Multimodal Intention Knowledge Distillation from Large Language Models for Social-Media Commonsense DiscoveryFeihong Lu, Weiqi Wang, Yangyifei Luo, Ziqin Zhu et al.ACM MM 2024 · 11 citations
