Can Language Models Laugh at YouTube Short-form Videos?
Dayoon Ko, Sangho Lee, Gunhee Kim
摘要
As short-form funny videos on social networks are gaining popularity, it becomes demanding for AI models to understand them for better communication with humans. Unfortunately, previous video humor datasets target specific domains such as speeches or sitcoms, and mostly focus on verbal cues. We curate a user-generated dataset of 10K multimodal funny videos from YouTube, called ExFunTube. Using a video filtering pipeline with GPT-3.5, we verify both verbal and visual elements contributing to humor. After filtering, we annotate each video with timestamps and text explanations for funny moments. Our ExFunTube is unique over existing datasets in that our videos cover a wide range of domains with various types of humor that necessitate a multimodal understanding of the content. Also, we develop a zero-shot video-to-text prompting to maximize video humor understanding of large language models (LLMs). With three different evaluation methods using automatic scores, rationale quality experiments, and human evaluations, we show that our prompting significantly improves LLMs’ ability for humor explanation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 被引用 6 次
- PunchBench: Benchmarking MLLMs in Multimodal Punchline ComprehensionKun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu 等ACL 2025 · 被引用 3 次
- v-HUB: A Benchmark for Video Humor Understanding from Vision and SoundZhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang 等ACL 2026
- SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about LaughterJungMok Lee, Sung-Bin Kim, Joohyun Chang, Lee Hyun 等ACL 2026
- SPIKE-RL: Video-LLMs meet Bayesian SurpriseSahithya Ravi, Aditya Chinchure, Raymond T. Ng, Leonid Sigal 等ICLR 2026
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
相关 Paper
- "I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou 等ACL 2026 · 被引用 1 次
- ExPUNations: Augmenting Puns with Keywords and ExplanationsJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone 等EMNLP 2022 · 被引用 7 次
- Talk Funny! A Large-Scale Humor Response Dataset with Chain-of-Humor InterpretationYuyan Chen, Yichen Yuan, Panjun Liu, Dayiheng Liu 等AAAI 2024 · 被引用 34 次
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language ModelsYifang Xu, Yunzhuo Sun, Benxiang Zhai, Ming Li 等AAAI 2025 · 被引用 17 次
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu 等ACM MM 2023 · 被引用 10 次
