Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, Yejin Choi
Abstract
Large neural networks can now generate jokes, but do they really "understand" humor? We challenge AI models with three tasks derived from the New Yorker Cartoon Caption Contest: matching a joke to a cartoon, identifying a winning caption, and explaining why a winning caption is funny. These tasks encapsulate progressively more sophisticated aspects of "understanding" a cartoon; key elements are the complex, often surprising relationships between images and captions and the frequent inclusion of indirect and playful allusions to human experience and culture. We investigate both multimodal and language-only models: the former are challenged with the cartoon images directly, while the latter are given multifaceted descriptions of the visual scene to simulate human-level visual understanding. We find that both types of models struggle at all three tasks. For example, our best multimodal models fall 30 accuracy points behind human performance on the matching task, and, even when provided ground-truth visual scene descriptors, human-authored explanations are preferred head-to-head over the best machine-authored ones (few-shot GPT-4) in more than 2/3 of cases. We release models, code, leaderboard, and corpus, which includes newly-gathered annotations describing the image's locations/entities, what's unusual in the scene, and an explanation of the joke. Human-authored: When drinking coffee or tea, people often add cream, and may ask others to pass it if it's on the other side of a table. But here, the mugs are huge, so instead of asking for a small cup of cream, they are asking for the entire cow, which is the appropriately-sized cream dispenser for these huge drinks. From Description (5-shot GPT 3.5): "Pass the cow" is an example of a non sequitur, something that looks like a logical thing to say, but doesn't make sense in context. The humor in this cartoon comes from the large size of the coffee mugs: they are so large that they resemble buckets rather than mugs, thus making the request to "pass the cow" almost reasonable. From Pixels (OFA + T5-11B): The joke is that the meeting participants are actually sitting on chairs made out of coffee mugs, which is an unlikely location for the discussion. Instead of asking for another mug of coffee, the person at the head of the table simply asks for "the cow", or a coffee machine. Explanation Generation Task 3 Can you please pass the cow? ❌ I'd kill for some cream cheese. vs. ✅ Can you please pass the cow? ❌ Welcome to Insomniacs Anonymous vs. 🏆 Can you please pass the cow? Matching Quality Ranking Explanation Generation You both know Jane A reference to Jane Goodall. Goodall is a scientist who is famous for studying chimpanzees, as represented by the ape at the party. This party is likely a scientific conference on biology, but the unusual part is that the subject of the study, the chimp, is invited. Both the peer scientist and the chimpanzee know Goodall, but for different reasons. A) I always figured hell would be less ironic. B) You both know Jane C) I'd better give it a little longer. It's a really tough case. D) And then I thought 'Wow, my cat really is kind of sexy.' E) We'll eventually miss him. 🏆 You both know Jane -vs-Accounting meet archives. Publicly, we are still saying there are no side effects This is a board meeting of a shady pharmaceutical company. The drug the company makes has the side effect of turning people into cartoon monsters, and most everyone at the company has taken it. Nonetheless, they are choosing not to warn the public. This plays upon a common belief that pharmaceutical companies care more about profits than they do the well-being of their patients. A) Can I interest you in an offshore account? B) So how much of the story is autobiographical? C) Don't give me that holier-than-thou attitude! D) They give me free drinks if I keep my tray table down.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers37
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu et al.ACL 2025 · 236 citations
- Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional ImagesNitzan Bitton Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt et al.ICCV 2023 · 92 citations
- Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMsYi Zhang, Bolin Ni, Xin-Sheng Chen, Hengrui Zhang et al.ICLR 2026 · 30 citations
- Multimodal Learning Without Labeled Multimodal Data: Guarantees and ApplicationsPaul Pu Liang, Chun Kai Ling, Yun Cheng, Alexander Obolenskiy et al.ICLR 2024 · 25 citations
- Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous ContradictionsZhe Hu, Tuo Liang, Jing Li, Yiren Lu et al.NeurIPS 2024 · 21 citations
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with LanguageAndy Zeng, Maria Attarian, Brian Ichter, Krzysztof Marcin Choromanski et al.ICLR 2023 · 171 citations
Related papers
- On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption GenerationWenbo Shang, Yuxi Sun, Jing Ma, Xin HuangICLR 2026 · 3 citations
- PunchBench: Benchmarking MLLMs in Multimodal Punchline ComprehensionKun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu et al.ACL 2025 · 3 citations
- HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor GenerationJiajun Zhang, Shijia Luo, Ruikang Zhang, Qi SuCVPR 2026 · 4 citations
- HumorDB: Can AI Understand Graphical Humor?Veedant Jain, Gabriel Kreiman, Felipe dos Santos Alves FeitosaICCV 2025 · 1 citation
- Can Language Models Laugh at YouTube Short-form Videos?Dayoon Ko, Sangho Lee, Gunhee KimEMNLP 2023 · 4 citations
