OxfordTVG-HIC: Can Machine Make Humorous Captions from Images?
Runjia Li, Shuyang Sun, Mohamed Elhoseiny, Philip H. S. Torr
Abstract
This paper presents OxfordTVG-HIC (Humorous Image Captions), a large-scale dataset for humour generation and understanding. Humour is an abstract, subjective, and context-dependent cognitive construct involving several cognitive factors, making it a challenging task to generate and interpret. Hence, humour generation and understanding can serve as a new task for evaluating the ability of deep-learning methods to process abstract and subjective information. Due to the scarcity of data, humourrelated generation tasks such as captioning remain underexplored. To address this gap, OxfordTVG-HIC offers approximately 2.9M image-text pairs with humour scores to train a generalizable humour captioning model. Contrary to existing captioning datasets, OxfordTVG-HIC features a wide range of emotional and semantic diversity resulting in out-of-context examples that are particularly conducive to generating humour. Moreover, OxfordTVG-HIC is curated devoid of offensive content. We also show how OxfordTVG-HIC can be leveraged for evaluating the humour of a generated text. Through explainability analysis of the trained models, we identify the visual and linguistic cues influential for evoking humour prediction (and generation). We observe qualitatively that these cues are aligned with the benign violation theory of humour in cognitive psychology.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- EmoVIT: Revolutionizing Emotion Insights with Visual Instruction TuningHongxia Xie, Chu-Jun Peng, Yu-Wen Tseng, Hung-Jen Chen et al.CVPR 2024 · 21 citations
- Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor GenerationShanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen et al.CVPR 2024 · 16 citations
- HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor GenerationJiajun Zhang, Shijia Luo, Ruikang Zhang, Qi SuCVPR 2026 · 4 citations
- HumorDB: Can AI Understand Graphical Humor?Veedant Jain, Gabriel Kreiman, Felipe dos Santos Alves FeitosaICCV 2025 · 1 citation
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Humor Knowledge Enriched Transformer for Understanding Multimodal HumorMd. Kamrul Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh et al.AAAI 2021 · 98 citations
Related papers
- ExPUNations: Augmenting Puns with Keywords and ExplanationsJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone et al.EMNLP 2022 · 7 citations
- On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption GenerationWenbo Shang, Yuxi Sun, Jing Ma, Xin HuangICLR 2026 · 3 citations
- "I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou et al.ACL 2026 · 1 citation
- v-HUB: A Benchmark for Video Humor Understanding from Vision and SoundZhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang et al.ACL 2026
- Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption ContestJack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee et al.ACL 2023 · 29 citations
