Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee, Youngjae Yu
摘要
Humans possess multimodal literacy, allowing them to actively integrate information from various modalities to form reasoning. Faced with challenges like lexical ambiguity in text, we supplement this with other modalities, such as thumbnail images or textbook illustrations. Is it possible for machines to achieve a similar multimodal understanding capability? In response, we present Understanding Pun with Image Explanations ( UNPIE) 1 , a novel benchmark designed to assess the impact of multimodal inputs in resolving lexical ambiguities. Puns serve as the ideal subject for this evaluation due to their intrinsic ambiguity. Our dataset includes 1,000 puns, each accompanied by an image that explains both meanings. We pose three multimodal challenges with the annotations to assess different aspects of multimodal literacy; Pun Grounding, Disambiguation, and Reconstruction. The results 2 indicate that various Socratic Models and Visual-Language Models improve over the text-only models when given visual context, particularly as the complexity of the tasks increases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- "I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou 等ACL 2026 · 被引用 1 次
- PunMemeCN: A Benchmark to Explore Vision-Language Models' Understanding of Chinese Pun MemesZhijun Xu, Siyu Yuan, Yiqiao Zhang, Jingyu Sun 等EMNLP 2025
- MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language ModelsXiaolong Wang, Zhaolu Kang, Wangyuxuan Zhai, Xinyue Lou 等EMNLP 2025
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- PunchBench: Benchmarking MLLMs in Multimodal Punchline ComprehensionKun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu 等ACL 2025 · 被引用 3 次
- VAGUE: Visual Contexts Clarify Ambiguous ExpressionsHeejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung 等ICCV 2025 · 被引用 1 次
- ExPUNations: Augmenting Puns with Keywords and ExplanationsJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone 等EMNLP 2022 · 被引用 7 次
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 被引用 6 次
- YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language ModelsAbhilash Nandy, Yash Agarwal, Ashish Patwa, Millon Madhur Das 等EMNLP 2024 · 被引用 2 次
