GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
Leonor Veloso, Hinrich Schütze
Abstract
Recent works have analyzed the impact of individual components of neural networks on gendered predictions, often with a focus on mitigating gender bias. However, mechanistic interpretations of gender tend to (i) focus on a very specific gender-related task, such as gendered pronoun prediction, or (ii) fail to distinguish between the production of factually gendered outputs (the correct assumption of gender given a word that carries gender as a semantic property) and gender biased outputs (based on a stereotype). To address these issues, we curate , a benchmark to assess gender knowledge and gender bias in language models across different types of gender-related predictions. allows us to identify and analyze circuits and individual neurons responsible for gendered predictions. We test the impact of neuron ablation on benchmarks for disentangling stereotypical and factual gender (DiFair and the test set of GKnow), as well as StereoSet. Results show that gender bias and factual gender are severely entangled on the level of both circuits and neurons, entailing that ablation is an unreliable debiasing method. Furthermore, we show that benchmarks for evaluating gender bias can hide the decrease in factual gender knowledge that accompanies neuron ablation. We curate GKnow as a contribution to the continuous development of robust gender bias benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbf454c1-394d-4671-a5ed-ae0307004720Builds on18
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell et al.NeurIPS 2023 · 305 citations
- Linearity of Relation Decoding in Transformer Language ModelsEvan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng et al.ICLR 2024 · 163 citations
- Knowledge Circuits in Pretrained TransformersYunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang et al.NeurIPS 2024 · 71 citations
- Journey to the Center of the Knowledge Neurons: Discoveries of Language-Independent Knowledge Neurons and Degenerate Knowledge NeuronsYuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu et al.AAAI 2024 · 64 citations
Related papers
- Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark DatasetsMahdi Zakizadeh, Mohammad Taher PilehvarEMNLP 2025
- Are Models Biased on Text without Gender-related Language?Catarina G. Belém, Preethi Seshadri, Yasaman Razeghi, Sameer SinghICLR 2024 · 16 citations
- Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language ModelsZhengyang Shan, Emily Diana, Jiawei ZhouACL 2025 · 3 citations
- Bias in Gender Bias Benchmarks: How Spurious Features Distort EvaluationYusuke Hirota, Ryo Hachiuma, Boyi Li, Ximing Lu et al.ICCV 2025
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
