Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
Francisco Vargas, Ryan Cotterell
摘要
Bolukbasi et al. (2016) presents one of the first gender bias mitigation techniques for word embeddings. Their method takes pre-trained word embeddings as input and attempts to isolate a linear subspace that captures most of the gender bias in the embeddings. As judged by an analogical evaluation task, their method virtually eliminates gender bias in the embeddings. However, an implicit and untested assumption of their method is that the bias subspace is actually linear. In this work, we generalize their method to a kernelized, non-linear version. We take inspiration from kernel principal component analysis and derive a nonlinear bias isolation technique. We discuss and overcome some of the practical drawbacks of our method for non-linear gender bias mitigation in word embeddings and analyze empirically whether the bias subspace is actually linear. Our analysis shows that gender bias is in fact well captured by a linear subspace, justifying the assumption of Bolukbasi et al. (2016) .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Linear Adversarial Concept ErasureShauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan CotterellICML 2022 · 被引用 89 次
- Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-TuningJulian Minder, Clément Dumas, Caden Juang, Bilal Chughtai 等NeurIPS 2025 · 被引用 32 次
- Assessing the Reliability of Word Embedding Gender Bias MeasuresYupei Du, Qixiang Fang, Dong NguyenEMNLP 2021 · 被引用 13 次
- What does a Text Classifier Learn about Morality? An Explainable Method for Cross-Domain Comparison of Moral RhetoricEnrico Liscio, Oscar Araque, Lorenzo Gatti, Ionut Constantinescu 等ACL 2023 · 被引用 12 次
- Emergence of Linear Truth Encodings in Language ModelsShauli Ravfogel, Gilad Yehudai, Tal Linzen, Joan Bruna 等NeurIPS 2025 · 被引用 12 次
相关 Paper
- Double-Hard Debias: Tailoring Word Embeddings for Gender Bias MitigationTianlu Wang, Xi Victoria Lin, Nazneen Fatema Rajani, Bryan McCann 等ACL 2020 · 被引用 42 次
- A Causal Inference Method for Reducing Gender Bias in Word Embedding RelationsZekun Yang, Juan FengAAAI 2020 · 被引用 40 次
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 被引用 195 次
- Towards Understanding Gender Bias in Relation ExtractionAndrew Gaut, Tony Sun, Shirlyn Tang, Yuxin Huang 等ACL 2020 · 被引用 7 次
- Adversarial Concept Erasure in Kernel SpaceShauli Ravfogel, Francisco Vargas, Yoav Goldberg, Ryan CotterellEMNLP 2022 · 被引用 11 次
