Multimodal Sarcasm Target Identification in Tweets
Jiquan Wang, Lin Sun, Yi Liu, Meizhi Shao, Zengwei Zheng
摘要
Sarcasm is important to sentiment analysis on social media. Sarcasm Target Identification (STI) deserves further study to understand sarcasm in depth. However, text lacking context or missing sarcasm target makes target identification very difficult. In this paper, we introduce multimodality to STI and present Multimodal Sarcasm Target Identification (MSTI) task. We propose a novel multi-scale cross-modality model that can simultaneously perform textual target labeling and visual target detection. In the model, we extract multi-scale visual features to enrich spatial information for different sized visual sarcasm targets. We design a set of convolution networks to unify multi-scale visual features with textual features for cross-modal attention learning, and correspondingly a set of transposed convolution networks to restore multi-scale visual information. The results show that visual clues can improve the performance of TSTI by a large margin, and VSTI achieves good accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- FigMemes: A Dataset for Figurative Language Identification in Politically-Opinionated MemesChen Liu, Gregor Geigle, Robin Krebs, Iryna GurevychEMNLP 2022 · 被引用 15 次
- CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal ModelsZixin Chen, Hongzhan Lin, Ziyang Luo, Mingfei Cheng 等ACL 2024 · 被引用 10 次
- DocMSU: A Comprehensive Benchmark for Document-Level Multimodal Sarcasm UnderstandingHang Du, Guoshun Nan, Sicheng Zhang, Binzhu Xie 等AAAI 2024 · 被引用 9 次
- Well, Now We Know! Unveiling Sarcasm: Initiating and Exploring Multimodal Conversations with ReasoningGopendra Vikram Singh, Mauajama Firdaus, Dushyant Singh Chauhan, Asif Ekbal 等AAAI 2024 · 被引用 6 次
- DIP: Dual Incongruity Perceiving Network for Sarcasm DetectionChangsong Wen, Guoli Jia, Jufeng YangCVPR 2023
它引用的顶会 Paper3
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationNan Xu, Zhixiong Zeng, Wenji MaoACL 2020 · 被引用 153 次
相关 Paper
- MSTI-Plus: Introducing Non-Sarcasm Reference Materials to Enhance Multimodal Sarcasm Target IdentificationFengmao Lv, Mengting Xiong, Junlin Fang, Lingli Zhang 等WWW 2025 · 被引用 1 次
- Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal GraphsBin Liang, Chenwei Lou, Xiang Li, Lin Gui 等ACM MM 2021 · 被引用 128 次
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen 等AAAI 2023 · 被引用 84 次
- Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional NetworkBin Liang, Chenwei Lou, Xiang Li, Min Yang 等ACL 2022 · 被引用 151 次
- Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge EnhancementHui Liu, Wenya Wang, Haoliang LiEMNLP 2022 · 被引用 91 次
