OxfordTVG-HIC: Can Machine Make Humorous Captions from Images?
Runjia Li, Shuyang Sun, Mohamed Elhoseiny, Philip H. S. Torr
摘要
This paper presents OxfordTVG-HIC (Humorous Image Captions), a large-scale dataset for humour generation and understanding. Humour is an abstract, subjective, and context-dependent cognitive construct involving several cognitive factors, making it a challenging task to generate and interpret. Hence, humour generation and understanding can serve as a new task for evaluating the ability of deep-learning methods to process abstract and subjective information. Due to the scarcity of data, humourrelated generation tasks such as captioning remain underexplored. To address this gap, OxfordTVG-HIC offers approximately 2.9M image-text pairs with humour scores to train a generalizable humour captioning model. Contrary to existing captioning datasets, OxfordTVG-HIC features a wide range of emotional and semantic diversity resulting in out-of-context examples that are particularly conducive to generating humour. Moreover, OxfordTVG-HIC is curated devoid of offensive content. We also show how OxfordTVG-HIC can be leveraged for evaluating the humour of a generated text. Through explainability analysis of the trained models, we identify the visual and linguistic cues influential for evoking humour prediction (and generation). We observe qualitatively that these cues are aligned with the benign violation theory of humour in cognitive psychology.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- EmoVIT: Revolutionizing Emotion Insights with Visual Instruction TuningHongxia Xie, Chu-Jun Peng, Yu-Wen Tseng, Hung-Jen Chen 等CVPR 2024 · 被引用 21 次
- Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor GenerationShanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen 等CVPR 2024 · 被引用 16 次
- HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor GenerationJiajun Zhang, Shijia Luo, Ruikang Zhang, Qi SuCVPR 2026 · 被引用 4 次
- HumorDB: Can AI Understand Graphical Humor?Veedant Jain, Gabriel Kreiman, Felipe dos Santos Alves FeitosaICCV 2025 · 被引用 1 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
- Humor Knowledge Enriched Transformer for Understanding Multimodal HumorMd. Kamrul Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh 等AAAI 2021 · 被引用 98 次
相关 Paper
- ExPUNations: Augmenting Puns with Keywords and ExplanationsJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone 等EMNLP 2022 · 被引用 7 次
- On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption GenerationWenbo Shang, Yuxi Sun, Jing Ma, Xin HuangICLR 2026 · 被引用 3 次
- "I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou 等ACL 2026 · 被引用 1 次
- v-HUB: A Benchmark for Video Humor Understanding from Vision and SoundZhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang 等ACL 2026
- Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption ContestJack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee 等ACL 2023 · 被引用 29 次
