Social-Group-Agnostic Bias Mitigation via the Stereotype Content Model
Ali Omrani, Alireza Salkhordeh Ziabari, Charles Yu, Preni Golazizian, Brendan Kennedy, Mohammad Atari, Heng Ji, Morteza Dehghani
摘要
Existing bias mitigation methods require socialgroup-specific word pairs (e.g., "man" -"woman") for each social attribute (e.g., gender), restricting the bias mitigation to only one specified social attribute. Further, this constraint renders such methods impractical and costly for mitigating bias in understudied and/or unmarked social groups. We propose that the Stereotype Content Model (SCM) -a theoretical framework developed in social psychology for understanding the content of stereotyping -can help debiasing efforts to become social-group-agnostic by capturing the underlying connection between bias and stereotypes. SCM proposes that the content of stereotypes map to two psychological dimensions of warmth and competence. Using only pairs of terms for these two dimensions (e.g., warmth: "genuine" -"fake"; competence: "smart" -"stupid"), we perform debiasing with established methods on both pretrained word embeddings and large language models. We demonstrate that our social-groupagnostic, SCM-based debiasing technique performs comparably to group-specific debiasing on multiple bias benchmarks, but has theoretical and practical advantages over existing approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- NORMSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-FlyYi Fung, Tuhin Chakrabarty, Hao Guo, Owen Rambow 等EMNLP 2023 · 被引用 17 次
- The Nature of NLP: Analyzing Contributions in NLP PapersAniket Pramanick, Yufang Hou, Saif M. Mohammad, Iryna GurevychACL 2025 · 被引用 9 次
- Word Embeddings Are Steers for Language ModelsChi Han, Jialiang Xu, Manling Li, Yi Fung 等ACL 2024 · 被引用 8 次
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius 等KDD 2024 · 被引用 7 次
- A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI EvaluationsAida Mostafazadeh Davani, Sunipa Dev, Héctor Pérez-Urbina, Vinodkumar PrabhakaranEMNLP 2025 · 被引用 6 次
它引用的顶会 Paper11
- Process for Adapting Language Models to Society (PALMS) with Values-Targeted DatasetsIrene Solaiman, Christy DennisonNeurIPS 2021 · 被引用 276 次
- Towards Debiasing Sentence RepresentationsPaul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim 等ACL 2020 · 被引用 149 次
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 被引用 94 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- ADEPT: A DEbiasing PrompT FrameworkKe Yang, Charles Yu, Yi Ren Fung, Manling Li 等AAAI 2023 · 被引用 40 次
相关 Paper
- Understanding and Countering Stereotypes: A Computational Approach to the Stereotype Content ModelKathleen C. Fraser, Isar Nejadgholi, Svetlana KiritchenkoACL 2021
- StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language ModelsSullam Jeoung, Yubin Ge, Jana DiesnerEMNLP 2023 · 被引用 3 次
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen 等ICLR 2026 · 被引用 6 次
- Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language ModelsYisong Xiao, Aishan Liu, Siyuan Liang, Xianglong Liu 等ISSTA 2025 · 被引用 2 次
- A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector SpacesAnne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan VulicAAAI 2020 · 被引用 68 次
