You Told Me That Joke Twice: A Systematic Investigation of Transferability and Robustness of Humor Detection Models
Alexander Baranov, Vladimir Kniazhevsky, Pavel Braslavski
Abstract
In this study, we focus on automatic humor detection, a highly relevant task for conversational AI. To date, there are several English datasets for this task, but little research on how models trained on them generalize and behave in the wild. To fill this gap, we carefully analyze existing datasets, train RoBERTabased and Naïve Bayes classifiers on each of them, and test on the rest. Training and testing on the same dataset yields good results, but the transferability of the models varies widely. Models trained on datasets with jokes from different sources show better transferability, while the amount of training data has a smaller impact. The behavior of the models on out-of-domain data is unstable, suggesting that some of the models overfit, while others learn non-specific humor characteristics. An adversarial attack shows that models trained on pun datasets are less robust. We also evaluate the sense of humor of the chatGPT and Flan-UL2 models in a zero-shot scenario. The LLMs demonstrate competitive results on humor datasets and a more stable behavior on out-of-domain data. We believe that the obtained results will facilitate the development of new datasets and evaluation methodologies in the field of computational humor. We've made all the data from the study and the trained models publicly available: https://github.com/ Humor-Research/Humor-detection . Tell me which are funny, which are not -and which get a giggle first time but are cold pancakes without honey to hear twice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12187668-4431-4f9a-bb69-11a2fc69fa92Cited by top-tier papers3
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 6 citations
- Assessing the Capabilities of LLMs in Humor: A Multi-dimensional Analysis of Oogiri Generation and EvaluationRitsu Sakabe, Hwichan Kim, Tosho Hirasawa, Mamoru KomachiAAAI 2026
- It's Not Bragging If You Can Back It Up: Can LLMs Understand Braggings?Jingjie Zeng, Huayang Li, Liang Yang, Yuanyuan Sun et al.ACL 2025
Builds on5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee et al.ICLR 2023 · 158 citations
- UL2: Unifying Language Learning ParadigmsYi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia et al.ICLR 2023 · 97 citations
- Humor Detection in Product Question Answering SystemsYftah Ziser, Elad Kravi, David CarmelSIGIR 2020 · 20 citations
Related papers
- "What do you call a dog that is incontrovertibly true? Dogma": Testing LLM Generalization through HumorAlessio Cocchieri, Luca Ragazzi, Paolo Italiani, Giuseppe Tagliavini et al.ACL 2025
- Talk Funny! A Large-Scale Humor Response Dataset with Chain-of-Humor InterpretationYuyan Chen, Yichen Yuan, Panjun Liu, Dayiheng Liu et al.AAAI 2024 · 34 citations
- Can Language Models Make Fun? A Case Study in Chinese Comical CrosstalkJianquan Li, Xiangbo Wu, Xiaokang Liu, Qianqian Xie et al.ACL 2023 · 2 citations
- HumorDB: Can AI Understand Graphical Humor?Veedant Jain, Gabriel Kreiman, Felipe dos Santos Alves FeitosaICCV 2025 · 1 citation
- "I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou et al.ACL 2026 · 1 citation
