S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm Detection
Zhihong Zhu, Fan Zhang, Yunyan Zhang, Jinghan Sun, Guimin Hu, Hao Wu, Yuyan Chen, Bowen Xing, Xian Wu
Abstract
Multimodal sarcasm detection (MSD) aims to identify sarcasm polarity from diverse modalities (i.e., image–text pairs), a task that has received increasing attention. While significant progress has been made, existing approaches still face two major issues: lack of explainability and weak generalizability. In this paper, we introduce a new large vision–language model (LVLM) dubbed S³-MSD for explainable and generalizable MSD through three key components. For explainability, we develop (1) a self-training paradigm that automatically bootstraps answers with explanations, and (2) a self-calibrating mechanism that rectifies flawed explanations. For generalizability, we design (3) a self-focusing module that amplifies visual semantic entities through preference optimization, thereby mitigating textual over-reliance. Experimental results on both in-distribution and out-of-distribution (OOD) benchmarks demonstrate that S³-MSD consistently outperforms state-of-the-art methods in detection performance. Furthermore, the proposed S³-MSD provides persuasive explanations, as verified by both quantitative metrics and human evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9fccf56c-c973-4c85-bff4-c04e16afbc47Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal ModelsZixin Chen, Hongzhan Lin, Ziyang Luo, Mingfei Cheng et al.ACL 2024 · 10 citations
- Nice Perfume. How Long Did You Marinate in It? Multimodal Sarcasm ExplanationPoorav Desai, Tanmoy Chakraborty, Md. Shad AkhtarAAAI 2022 · 49 citations
- MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm DetectionHaochen Zhao, Yuyao Kong, Yongxiu Xu, Gaopeng Gou et al.CVPR 2026 · 4 citations
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen et al.AAAI 2023 · 84 citations
- Debiasing Multimodal Sarcasm Detection with Contrastive LearningMengzhao Jia, Can Xie, Liqiang JingAAAI 2024 · 51 citations
