Debiasing Pretrained Text Encoders by Paying Attention to Paying Attention
Yacine Gaci, Boualem Benatallah, Fabio Casati, Khalid Benabdeslem
摘要
Natural Language Processing (NLP) models are found to exhibit discriminatory stereotypes across many social constructs, e.g. gender and race. In comparison to the progress made in reducing bias from static word embeddings, fairness in sentence-level text encoders received little consideration despite their wider applicability in contemporary NLP tasks. In this paper, we propose a debiasing method for pre-trained text encoders that both reduces social stereotypes, and inflicts next to no semantic damage. Unlike previous studies that directly manipulate the embeddings, we suggest to dive deeper into the operation of these encoders, and pay more attention to the way they pay attention to different social groups. We find that stereotypes are also encoded in the attention layer. Then, we work on model debiasing by redistributing the attention scores of a text encoder such that it forgets any preference to historically advantaged groups, and attends to all social classes with the same intensity. Our experiments confirm that reducing bias from attention effectively mitigates it from the model’s text representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LLM Bias Detection and Mitigation through the Lens of Desired DistributionsIngroj Shrestha, Padmini SrinivasanEMNLP 2025 · 被引用 3 次
- KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language ModelsSeorin Kim, Dongyoung Lee, Jaejin LeeEMNLP 2025
它引用的顶会 Paper13
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 被引用 195 次
- An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language ModelsNicholas Meade, Elinor Poole-Dayan, Siva ReddyACL 2022 · 被引用 160 次
- Towards Debiasing Sentence RepresentationsPaul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim 等ACL 2020 · 被引用 149 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
相关 Paper
- FairFil: Contrastive Neural Debiasing Method for Pretrained Text EncodersPengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si 等ICLR 2021 · 被引用 50 次
- Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningFan Zhou, Yuzhou Mao, Liu Yu, Yi Yang 等ACL 2023 · 被引用 21 次
- Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsYue Guo, Yi Yang, Ahmed AbbasiACL 2022
- Mitigating Biases in Language Models via Bias UnlearningDianqing Liu, Yi Liu, Guoqing Jin, Zhendong MaoEMNLP 2025 · 被引用 4 次
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen 等ICLR 2026 · 被引用 6 次
