MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
Haochen Zhao, Yuyao Kong, Yongxiu Xu, Gaopeng Gou, Hongbo Xu, Yubin Wang, Haoliang Zhang
Abstract
Despite progress in multimodal sarcasm detection, existing datasets and methods predominantly focus on single-image scenarios, overlooking potential semantic and affective relations across multiple images. This leaves a gap in modeling cases where sarcasm is triggered by multi-image cues in real-world settings. To bridge this gap, we introduce MMSD3.0, a new benchmark composed entirely of multi-image samples curated from tweets and Amazon reviews. We further propose the Cross-Image Reasoning Model (CIRM), which performs targeted cross-image sequence modeling to capture latent inter-image connections. In addition, we introduce a relevance-guided, fine-grained cross-modal fusion mechanism based on text-image correspondence to reduce information loss during integration. We establish a comprehensive suite of strong and representative baselines and conduct extensive experiments, showing that MMSD3.0 is an effective and reliable benchmark that better reflects real-world conditions. Moreover, CIRM demonstrates state-of-the-art performance across MMSD, MMSD2.0 and MMSD3.0, validating its effectiveness in both single-image and multi-image scenarios. Dataset and code are publicly available at https://github.com/ZHCMOONWIND/MMSD3.0.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationNan Xu, Zhixiong Zeng, Wenji MaoACL 2020 · 153 citations
- Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge EnhancementHui Liu, Wenya Wang, Haoliang LiEMNLP 2022 · 91 citations
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen et al.AAAI 2023 · 84 citations
- MoBA: Mixture of Bi-directional Adapter for Multi-modal Sarcasm DetectionYifeng Xie, Zhihong Zhu, Xin Chen, Zhanpeng Chen et al.ACM MM 2024 · 11 citations
Related papers
- Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal GraphsBin Liang, Chenwei Lou, Xiang Li, Lin Gui et al.ACM MM 2021 · 128 citations
- DocMSU: A Comprehensive Benchmark for Document-Level Multimodal Sarcasm UnderstandingHang Du, Guoshun Nan, Sicheng Zhang, Binzhu Xie et al.AAAI 2024 · 9 citations
- Well, Now We Know! Unveiling Sarcasm: Initiating and Exploring Multimodal Conversations with ReasoningGopendra Vikram Singh, Mauajama Firdaus, Dushyant Singh Chauhan, Asif Ekbal et al.AAAI 2024 · 6 citations
- DIP: Dual Incongruity Perceiving Network for Sarcasm DetectionChangsong Wen, Guoli Jia, Jufeng YangCVPR 2023
- Multimodal Sarcasm Target Identification in TweetsJiquan Wang, Lin Sun, Yi Liu, Meizhi Shao et al.ACL 2022 · 28 citations
