CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal Models
Zixin Chen, Hongzhan Lin, Ziyang Luo, Mingfei Cheng, Jing Ma, Guang Chen
摘要
Social media abounds with multimodal sarcasm, and identifying sarcasm targets is particularly challenging due to the implicit incongruity not directly evident in the text and image modalities. Current methods for Multimodal Sarcasm Target Identification (MSTI) predominantly focus on superficial indicators in an end-to-end manner, overlooking the nuanced understanding of multimodal sarcasm conveyed through both the text and image. This paper proposes a versatile MSTI framework with a coarse-tofine paradigm, by augmenting sarcasm explainability with reasoning and pre-training knowledge. Inspired by the powerful capacity of Large Multimodal Models (LMMs) on multimodal reasoning, we first engage LMMs to generate competing rationales for coarser-grained pre-training of a small language model on multimodal sarcasm detection. We then propose fine-tuning the model for finer-grained sarcasm target identification. Our framework is thus empowered to adeptly unveil the intricate targets within multimodal sarcasm and mitigate the negative impact posed by potential noise inherently in LMMs. Experimental results demonstrate that our model far outperforms state-ofthe-art MSTI methods, and markedly exhibits explainability in deciphering sarcasm as well.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful ReasoningTianrun Xu, Haoda Jing, Ye Li, Yuquan Wei 等ICML 2026 · 被引用 8 次
- On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMsHerun Wan, Minnan Luo, Zhixiong Su, Guang Dai 等ACL 2025 · 被引用 5 次
- MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm DetectionHaochen Zhao, Yuyao Kong, Yongxiu Xu, Gaopeng Gou 等CVPR 2026 · 被引用 4 次
- MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language ModelsZixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo 等EMNLP 2025
- MuVaC: A Variational Causal Framework for Multimodal Sarcasm Understanding in DialoguesDiandian Guo, Fangfang Yuan, Cong Cao, Xixun Lin 等WWW 2026
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
相关 Paper
- Multimodal Sarcasm Target Identification in TweetsJiquan Wang, Lin Sun, Yi Liu, Meizhi Shao 等ACL 2022 · 被引用 28 次
- MSTI-Plus: Introducing Non-Sarcasm Reference Materials to Enhance Multimodal Sarcasm Target IdentificationFengmao Lv, Mengting Xiong, Junlin Fang, Lingli Zhang 等WWW 2025 · 被引用 1 次
- S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm DetectionZhihong Zhu, Fan Zhang, Yunyan Zhang, Jinghan Sun 等AAAI 2026
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen 等AAAI 2023 · 被引用 84 次
- Miko: Multimodal Intention Knowledge Distillation from Large Language Models for Social-Media Commonsense DiscoveryFeihong Lu, Weiqi Wang, Yangyifei Luo, Ziqin Zhu 等ACM MM 2024 · 被引用 11 次
