"I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?
Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou, Yuyuan Li, Tianyu Du, Jun Wang, Zhihui Fu, Jinbao Li, Shouling Ji
摘要
Puns are a common form of rhetorical wordplay that exploits polysemy and phonetic similarity to create humor. In multimodal puns, visual and textual elements synergize to ground the literal sense and evoke the figurative meaning simultaneously. Although Vision-Language Models (VLMs) are widely used in multimodal understanding and generation, their ability to understand puns has not been systematically studied due to a scarcity of rigorous benchmarks. To address this, we first propose a multimodal pun generation pipeline. We then introduce MultiPun, a dataset comprising diverse types of puns alongside adversarial non-pun distractors. Our evaluation reveals that most models struggle to distinguish genuine puns from these distractors. Moreover, we propose both prompt-level and model-level strategies to enhance pun comprehension, with an average improvement of 16.5% in F1 scores. Our findings provide valuable insights for developing future VLMs that master the subtleties of human-like humor via cross-modal reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal GuardrailsShuo Shi, Rui Yin, Naen Xu, Jiahao Chen 等KDD 2026 · 被引用 1 次
- Don't Reinvent the Wheel, Just Realign the Spokes: Resource-Efficient Federated Fine-Tuning via Rank-Wise Expert AssemblyYebo Wu, Jingguang Li, Zhijiang Guo, Li LiICML 2026
它引用的顶会 Paper15
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalZixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu 等AAAI 2025 · 被引用 59 次
- Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous ContradictionsZhe Hu, Tuo Liang, Jing Li, Yiren Lu 等NeurIPS 2024 · 被引用 21 次
- Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee 等EMNLP 2024 · 被引用 8 次
- HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Shiqi Zhang 等AAAI 2026 · 被引用 8 次
相关 Paper
- Pun Unintended: LLMs and the Illusion of Humor UnderstandingAlessandro Zangari, Matteo Marcuzzo, Andrea Albarelli, Mohammad Taher Pilehvar 等EMNLP 2025 · 被引用 1 次
- PunMemeCN: A Benchmark to Explore Vision-Language Models' Understanding of Chinese Pun MemesZhijun Xu, Siyu Yuan, Yiqiao Zhang, Jingyu Sun 等EMNLP 2025
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 被引用 6 次
- ExPUNations: Augmenting Puns with Keywords and ExplanationsJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone 等EMNLP 2022 · 被引用 7 次
- "The Boating Store Had Its Best Sail Ever": Pronunciation-attentive Contextualized Pun RecognitionYichao Zhou, Jyun-Yu Jiang, Jieyu Zhao, Kai-Wei Chang 等ACL 2020 · 被引用 7 次
