Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, Yejin Choi
摘要
Large neural networks can now generate jokes, but do they really "understand" humor? We challenge AI models with three tasks derived from the New Yorker Cartoon Caption Contest: matching a joke to a cartoon, identifying a winning caption, and explaining why a winning caption is funny. These tasks encapsulate progressively more sophisticated aspects of "understanding" a cartoon; key elements are the complex, often surprising relationships between images and captions and the frequent inclusion of indirect and playful allusions to human experience and culture. We investigate both multimodal and language-only models: the former are challenged with the cartoon images directly, while the latter are given multifaceted descriptions of the visual scene to simulate human-level visual understanding. We find that both types of models struggle at all three tasks. For example, our best multimodal models fall 30 accuracy points behind human performance on the matching task, and, even when provided ground-truth visual scene descriptors, human-authored explanations are preferred head-to-head over the best machine-authored ones (few-shot GPT-4) in more than 2/3 of cases. We release models, code, leaderboard, and corpus, which includes newly-gathered annotations describing the image's locations/entities, what's unusual in the scene, and an explanation of the joke. Human-authored: When drinking coffee or tea, people often add cream, and may ask others to pass it if it's on the other side of a table. But here, the mugs are huge, so instead of asking for a small cup of cream, they are asking for the entire cow, which is the appropriately-sized cream dispenser for these huge drinks. From Description (5-shot GPT 3.5): "Pass the cow" is an example of a non sequitur, something that looks like a logical thing to say, but doesn't make sense in context. The humor in this cartoon comes from the large size of the coffee mugs: they are so large that they resemble buckets rather than mugs, thus making the request to "pass the cow" almost reasonable. From Pixels (OFA + T5-11B): The joke is that the meeting participants are actually sitting on chairs made out of coffee mugs, which is an unlikely location for the discussion. Instead of asking for another mug of coffee, the person at the head of the table simply asks for "the cow", or a coffee machine. Explanation Generation Task 3 Can you please pass the cow? ❌ I'd kill for some cream cheese. vs. ✅ Can you please pass the cow? ❌ Welcome to Insomniacs Anonymous vs. 🏆 Can you please pass the cow? Matching Quality Ranking Explanation Generation You both know Jane A reference to Jane Goodall. Goodall is a scientist who is famous for studying chimpanzees, as represented by the ape at the party. This party is likely a scientific conference on biology, but the unusual part is that the subject of the study, the chimp, is invited. Both the peer scientist and the chimpanzee know Goodall, but for different reasons. A) I always figured hell would be less ironic. B) You both know Jane C) I'd better give it a little longer. It's a really tough case. D) And then I thought 'Wow, my cat really is kind of sexy.' E) We'll eventually miss him. 🏆 You both know Jane -vs-Accounting meet archives. Publicly, we are still saying there are no side effects This is a board meeting of a shady pharmaceutical company. The drug the company makes has the side effect of turning people into cartoon monsters, and most everyone at the company has taken it. Nonetheless, they are choosing not to warn the public. This plays upon a common belief that pharmaceutical companies care more about profits than they do the well-being of their patients. A) Can I interest you in an offshore account? B) So how much of the story is autobiographical? C) Don't give me that holier-than-thou attitude! D) They give me free drinks if I keep my tray table down.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu 等ACL 2025 · 被引用 236 次
- Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional ImagesNitzan Bitton Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt 等ICCV 2023 · 被引用 92 次
- Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMsYi Zhang, Bolin Ni, Xin-Sheng Chen, Hengrui Zhang 等ICLR 2026 · 被引用 30 次
- Multimodal Learning Without Labeled Multimodal Data: Guarantees and ApplicationsPaul Pu Liang, Chun Kai Ling, Yun Cheng, Alexander Obolenskiy 等ICLR 2024 · 被引用 25 次
- Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous ContradictionsZhe Hu, Tuo Liang, Jing Li, Yiren Lu 等NeurIPS 2024 · 被引用 21 次
它引用的顶会 Paper6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with LanguageAndy Zeng, Maria Attarian, Brian Ichter, Krzysztof Marcin Choromanski 等ICLR 2023 · 被引用 171 次
相关 Paper
- On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption GenerationWenbo Shang, Yuxi Sun, Jing Ma, Xin HuangICLR 2026 · 被引用 3 次
- PunchBench: Benchmarking MLLMs in Multimodal Punchline ComprehensionKun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu 等ACL 2025 · 被引用 3 次
- HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor GenerationJiajun Zhang, Shijia Luo, Ruikang Zhang, Qi SuCVPR 2026 · 被引用 4 次
- HumorDB: Can AI Understand Graphical Humor?Veedant Jain, Gabriel Kreiman, Felipe dos Santos Alves FeitosaICCV 2025 · 被引用 1 次
- Can Language Models Laugh at YouTube Short-form Videos?Dayoon Ko, Sangho Lee, Gunhee KimEMNLP 2023 · 被引用 4 次
