What Do You MEME? Generating Explanations for Visual Semantic Role Labelling in Memes
Shivam Sharma, Siddhant Agarwal, Tharun Suresh, Preslav Nakov, Md. Shad Akhtar, Tanmoy Chakraborty
摘要
Memes are powerful means for effective communication on social media. Their effortless amalgamation of viral visuals and compelling messages can have far-reaching implications with proper marketing. Previous research on memes has primarily focused on characterizing their affective spectrum and detecting whether the meme's message insinuates any intended harm, such as hate, offense, racism, etc. However, memes often use abstraction, which can be elusive. Here, we introduce a novel task - EXCLAIM, generating explanations for visual semantic role labeling in memes. To this end, we curate ExHVV, a novel dataset that offers natural language explanations of connotative roles for three types of entities - heroes, villains, and victims, encompassing 4,680 entities present in 3K memes. We also benchmark ExHVV with several strong unimodal and multimodal baselines. Moreover, we posit LUMEN, a novel multimodal, multi-task learning framework that endeavors to address EXCLAIM optimally by jointly learning to predict the correct semantic roles and correspondingly to generate suitable natural language explanations. LUMEN distinctly outperforms the best baseline across 18 standard natural language generation evaluation metrics. Our systematic evaluation and analyses demonstrate that characteristic multimodal cues required for adjudicating semantic roles are also helpful for generating suitable explanations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion RecognitionJian Chen, Yuxuan Hu, Haifeng Lu, Wei Wang 等ACM MM 2025 · 被引用 5 次
- Computational Meme Understanding: A SurveyKhoi P. N. Nguyen, Vincent NgEMNLP 2024 · 被引用 3 次
- MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language ModelsZhengyi Zhao, Shubo Zhang, Yuxi Zhang, Yanxi Zhao 等EMNLP 2025
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- Ask, Acquire, Understand: A Multimodal Agent-based Framework for Social Abuse Detection in MemesXuanrui Lin, Chao Jia, Junhui Ji, Hui Han 等WWW 2025 · 被引用 9 次
- MemeIntel: Explainable Detection of Propagandistic and Hateful MemesMohamed Bayan Kmainasi, Abul Hasnat, Md. Arid Hasan, Ali Ezzat Shahroor 等EMNLP 2025 · 被引用 1 次
- MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched ContextualizationShivam Sharma, Ramaneswaran S., Udit Arora, Md. Shad Akhtar 等ACL 2023 · 被引用 2 次
- MemeCap: A Dataset for Captioning and Interpreting MemesEunjeong Hwang, Vered ShwartzEMNLP 2023 · 被引用 14 次
- Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme ExplanationYubo Xie, Chenkai Wang, Zongyang Ma, Fahui MiaoEMNLP 2025
