You Told Me That Joke Twice: A Systematic Investigation of Transferability and Robustness of Humor Detection Models
Alexander Baranov, Vladimir Kniazhevsky, Pavel Braslavski
摘要
In this study, we focus on automatic humor detection, a highly relevant task for conversational AI. To date, there are several English datasets for this task, but little research on how models trained on them generalize and behave in the wild. To fill this gap, we carefully analyze existing datasets, train RoBERTabased and Naïve Bayes classifiers on each of them, and test on the rest. Training and testing on the same dataset yields good results, but the transferability of the models varies widely. Models trained on datasets with jokes from different sources show better transferability, while the amount of training data has a smaller impact. The behavior of the models on out-of-domain data is unstable, suggesting that some of the models overfit, while others learn non-specific humor characteristics. An adversarial attack shows that models trained on pun datasets are less robust. We also evaluate the sense of humor of the chatGPT and Flan-UL2 models in a zero-shot scenario. The LLMs demonstrate competitive results on humor datasets and a more stable behavior on out-of-domain data. We believe that the obtained results will facilitate the development of new datasets and evaluation methodologies in the field of computational humor. We've made all the data from the study and the trained models publicly available: https://github.com/ Humor-Research/Humor-detection . Tell me which are funny, which are not -and which get a giggle first time but are cold pancakes without honey to hear twice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 被引用 6 次
- Assessing the Capabilities of LLMs in Humor: A Multi-dimensional Analysis of Oogiri Generation and EvaluationRitsu Sakabe, Hwichan Kim, Tosho Hirasawa, Mamoru KomachiAAAI 2026
- It's Not Bragging If You Can Back It Up: Can LLMs Understand Braggings?Jingjie Zeng, Huayang Li, Liang Yang, Yuanyuan Sun 等ACL 2025
它引用的顶会 Paper5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- UL2: Unifying Language Learning ParadigmsYi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia 等ICLR 2023 · 被引用 97 次
- Humor Detection in Product Question Answering SystemsYftah Ziser, Elad Kravi, David CarmelSIGIR 2020 · 被引用 20 次
相关 Paper
- "What do you call a dog that is incontrovertibly true? Dogma": Testing LLM Generalization through HumorAlessio Cocchieri, Luca Ragazzi, Paolo Italiani, Giuseppe Tagliavini 等ACL 2025
- Talk Funny! A Large-Scale Humor Response Dataset with Chain-of-Humor InterpretationYuyan Chen, Yichen Yuan, Panjun Liu, Dayiheng Liu 等AAAI 2024 · 被引用 34 次
- Can Language Models Make Fun? A Case Study in Chinese Comical CrosstalkJianquan Li, Xiangbo Wu, Xiaokang Liu, Qianqian Xie 等ACL 2023 · 被引用 2 次
- HumorDB: Can AI Understand Graphical Humor?Veedant Jain, Gabriel Kreiman, Felipe dos Santos Alves FeitosaICCV 2025 · 被引用 1 次
- "I See What You Did There": Can Large Vision-Language Models Understand Multimodal Puns?Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou 等ACL 2026 · 被引用 1 次
