Investigating Counterfactual Unfairness in LLMs towards Identities through Humor
Shubin Kim, Yejin Son, Junyeong Park, Keummin Ka, Seungbeen Lee, Jaeyoung Lee, Hyeju Jang, Alice Oh, Youngjae Yu
摘要
Warning: This paper contains content that may be offensive or upsetting. Humor holds up a mirror to social perception: what we find funny often reflects who we are and how we judge others. When language models engage with humor, their reactions expose the social assumptions they have internalized from training data. In this paper, we investigate counterfactual unfairness through humor by observing how the model's responses change when we swap who speaks and who is addressed while holding other factors constant. Our framework spans three tasks: humor generation refusal, speaker intention inference, and relational/societal impact prediction, covering both identity-agnostic humor and identity-specific disparagement humor. We introduce interpretable bias metrics that capture asymmetric patterns under identity swaps. Experiments across state-of-the-art models reveal consistent relational disparities: jokes told by privileged speakers are refused up to 67.5% more often, judged as malicious 64.7% more frequently, and rated up to 1.5 points higher in social harm on a 5-point scale. These patterns highlight how sensitivity and stereotyping coexist in generative models, complicating efforts toward fairness and cultural alignment. 1 * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani 等EMNLP 2022 · 被引用 56 次
- ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in ContextVictoria R. Li, Yida Chen, Naomi SaphraEMNLP 2024 · 被引用 5 次
- Certifying Counterfactual Bias in LLMsIsha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi 等ICLR 2025 · 被引用 3 次
相关 Paper
- Assessing the Capabilities of LLMs in Humor: A Multi-dimensional Analysis of Oogiri Generation and EvaluationRitsu Sakabe, Hwichan Kim, Tosho Hirasawa, Mamoru KomachiAAAI 2026
- VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language ModelsChahat Raj, Bowen Wei, Aylin Caliskan, Antonios Anastasopoulos 等ACL 2026 · 被引用 3 次
- Talk Funny! A Large-Scale Humor Response Dataset with Chain-of-Humor InterpretationYuyan Chen, Yichen Yuan, Panjun Liu, Dayiheng Liu 等AAAI 2024 · 被引用 34 次
- Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language ModelsRyan Steed, Swetasudha Panda, Ari Kobren, Michael L. WickACL 2022 · 被引用 52 次
- Uncovering and Quantifying Social Biases in Code GenerationYan Liu, Xiaokang Chen, Yan Gao, Zhe Su 等NeurIPS 2023 · 被引用 47 次
