Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
Francisco Vargas, Ryan Cotterell
Abstract
Bolukbasi et al. (2016) presents one of the first gender bias mitigation techniques for word embeddings. Their method takes pre-trained word embeddings as input and attempts to isolate a linear subspace that captures most of the gender bias in the embeddings. As judged by an analogical evaluation task, their method virtually eliminates gender bias in the embeddings. However, an implicit and untested assumption of their method is that the bias subspace is actually linear. In this work, we generalize their method to a kernelized, non-linear version. We take inspiration from kernel principal component analysis and derive a nonlinear bias isolation technique. We discuss and overcome some of the practical drawbacks of our method for non-linear gender bias mitigation in word embeddings and analyze empirically whether the bias subspace is actually linear. Our analysis shows that gender bias is in fact well captured by a linear subspace, justifying the assumption of Bolukbasi et al. (2016) .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccc5f626-8f65-4441-96c7-ed0c0df5b1deCited by top-tier papers12
- Linear Adversarial Concept ErasureShauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan CotterellICML 2022 · 89 citations
- Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-TuningJulian Minder, Clément Dumas, Caden Juang, Bilal Chughtai et al.NeurIPS 2025 · 32 citations
- Assessing the Reliability of Word Embedding Gender Bias MeasuresYupei Du, Qixiang Fang, Dong NguyenEMNLP 2021 · 13 citations
- What does a Text Classifier Learn about Morality? An Explainable Method for Cross-Domain Comparison of Moral RhetoricEnrico Liscio, Oscar Araque, Lorenzo Gatti, Ionut Constantinescu et al.ACL 2023 · 12 citations
- Emergence of Linear Truth Encodings in Language ModelsShauli Ravfogel, Gilad Yehudai, Tal Linzen, Joan Bruna et al.NeurIPS 2025 · 12 citations
Related papers
- Double-Hard Debias: Tailoring Word Embeddings for Gender Bias MitigationTianlu Wang, Xi Victoria Lin, Nazneen Fatema Rajani, Bryan McCann et al.ACL 2020 · 42 citations
- A Causal Inference Method for Reducing Gender Bias in Word Embedding RelationsZekun Yang, Juan FengAAAI 2020 · 40 citations
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 195 citations
- Towards Understanding Gender Bias in Relation ExtractionAndrew Gaut, Tony Sun, Shirlyn Tang, Yuxin Huang et al.ACL 2020 · 7 citations
- Adversarial Concept Erasure in Kernel SpaceShauli Ravfogel, Francisco Vargas, Yoav Goldberg, Ryan CotterellEMNLP 2022 · 11 citations
