"I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?
Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou, Yuyuan Li, Tianyu Du, Jun Wang, Zhihui Fu, Jinbao Li, Shouling Ji
Abstract
Puns are a common form of rhetorical wordplay that exploits polysemy and phonetic similarity to create humor. In multimodal puns, visual and textual elements synergize to ground the literal sense and evoke the figurative meaning simultaneously. Although Vision-Language Models (VLMs) are widely used in multimodal understanding and generation, their ability to understand puns has not been systematically studied due to a scarcity of rigorous benchmarks. To address this, we first propose a multimodal pun generation pipeline. We then introduce MultiPun, a dataset comprising diverse types of puns alongside adversarial non-pun distractors. Our evaluation reveals that most models struggle to distinguish genuine puns from these distractors. Moreover, we propose both prompt-level and model-level strategies to enhance pun comprehension, with an average improvement of 16.5% in F1 scores. Our findings provide valuable insights for developing future VLMs that master the subtleties of human-like humor via cross-modal reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e076d41c-1277-41b7-88bc-42071e7a51b8Cited by top-tier papers2
- The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal GuardrailsShuo Shi, Rui Yin, Naen Xu, Jiahao Chen et al.KDD 2026 · 1 citation
- Don't Reinvent the Wheel, Just Realign the Spokes: Resource-Efficient Federated Fine-Tuning via Rank-Wise Expert AssemblyYebo Wu, Jingguang Li, Zhijiang Guo, Li LiICML 2026
Builds on15
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalZixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu et al.AAAI 2025 · 59 citations
- Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous ContradictionsZhe Hu, Tuo Liang, Jing Li, Yiren Lu et al.NeurIPS 2024 · 21 citations
- Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee et al.EMNLP 2024 · 8 citations
- HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Shiqi Zhang et al.AAAI 2026 · 8 citations
Related papers
- Pun Unintended: LLMs and the Illusion of Humor UnderstandingAlessandro Zangari, Matteo Marcuzzo, Andrea Albarelli, Mohammad Taher Pilehvar et al.EMNLP 2025 · 1 citation
- PunMemeCN: A Benchmark to Explore Vision-Language Models' Understanding of Chinese Pun MemesZhijun Xu, Siyu Yuan, Yiqiao Zhang, Jingyu Sun et al.EMNLP 2025
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 6 citations
- ExPUNations: Augmenting Puns with Keywords and ExplanationsJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone et al.EMNLP 2022 · 7 citations
- "The Boating Store Had Its Best Sail Ever": Pronunciation-attentive Contextualized Pun RecognitionYichao Zhou, Jyun-Yu Jiang, Jieyu Zhao, Kai-Wei Chang et al.ACL 2020 · 7 citations
