Humor Knowledge Enriched Transformer for Understanding Multimodal Humor
Md. Kamrul Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh, Rada Mihalcea, Louis-Philippe Morency, Ehsan Hoque
Abstract
Recognizing humor from a video utterance requires understanding the verbal and non-verbal components as well as incorporating the appropriate context and external knowledge. In this paper, we propose Humor Knowledge enriched Transformer (HKT) that can capture the gist of a multimodal humorous expression by integrating the preceding context and external knowledge. We incorporate humor centric external knowledge into the model by capturing the ambiguity and sentiment present in the language. We encode all the language, acoustic, vision, and humor centric features separately using Transformer based encoders, followed by a cross attention layer to exchange information among them. Our model achieves 77.36% and 79.41% accuracy in humorous punchline detection on UR-FUNNY and MUStaRD datasets -- achieving a new state-of-the-art on both datasets with the margin of 4.93% and 2.94% respectively. Furthermore, we demonstrate that our model can capture interpretable, humor-inducing patterns from all modalities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf0d6a78-db0c-4373-b645-93864b437030Cited by top-tier papers14
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou et al.ACM MM 2022 · 66 citations
- When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party DialoguesShivani Kumar, Atharva Kulkarni, Md. Shad Akhtar, Tanmoy ChakrabortyACL 2022 · 54 citations
- Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption ContestJack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee et al.ACL 2023 · 29 citations
- Is Sarcasm Detection a Step-by-Step Reasoning Process in Large Language Models?Ben Yao, Yazhou Zhang, Qiuchi Li, Jing QinAAAI 2025 · 29 citations
- Multimodal Learning Without Labeled Multimodal Data: Guarantees and ApplicationsPaul Pu Liang, Chun Kai Ling, Yun Cheng, Alexander Obolenskiy et al.ICLR 2024 · 25 citations
Builds on4
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
Related papers
- Can Language Models Laugh at YouTube Short-form Videos?Dayoon Ko, Sangho Lee, Gunhee KimEMNLP 2023 · 4 citations
- v-HUB: A Benchmark for Video Humor Understanding from Vision and SoundZhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang et al.ACL 2026
- "I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou et al.ACL 2026 · 1 citation
- Multiple Knowledge Syncretic Transformer for Natural Dialogue GenerationXiangyu Zhao, Longbiao Wang, Ruifang He, Ting Yang et al.WWW 2020 · 27 citations
- MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction ExpertsHaofei Yu, Zhengyang Qi, Lawrence Jang, Russ Salakhutdinov et al.EMNLP 2024 · 11 citations
