Investigating Gender Bias in Language Models Using Causal Mediation Analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, Stuart M. Shieber
Abstract
Many interpretation methods for neural models in natural language processing investigate how information is encoded inside hidden representations. However, these methods can only measure whether the information exists, not whether it is actually used by the model. We propose a methodology grounded in the theory of causal mediation analysis for interpreting which parts of a model are causally implicated in its behavior. The approach enables us to analyze the mechanisms that facilitate the flow of information from input to output through various model components, known as mediators. As a case study, we apply this methodology to analyzing gender bias in pre-trained Transformer language models. We study the role of individual neurons and attention heads in mediating gender bias across three datasets designed to gauge a model's sensitivity to gender bias. Our mediation analysis reveals that gender bias effects are concentrated in specific components of the model that may exhibit highly specialized behavior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d137eda-44f1-4c2e-9a64-c39bd461de09Cited by top-tier papers199
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 651 citations
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong et al.ICLR 2021 · 357 citations
Builds on4
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- A Causal Inference Method for Reducing Gender Bias in Word Embedding RelationsZekun Yang, Juan FengAAAI 2020 · 40 citations
- Toward Gender-Inclusive Coreference ResolutionYang Trista Cao, Hal Daumé IIIACL 2020 · 20 citations
Related papers
- Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation PerspectiveZhaotian Weng, Zijun Gao, Jerone Theodore Alexander Andrews, Jieyu ZhaoEMNLP 2024 · 3 citations
- Debiasing Algorithm through Model AdaptationTomasz Limisiewicz, David Marecek, Tomás MusilICLR 2024 · 24 citations
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisAlessandro Stolfo, Yonatan Belinkov, Mrinmaya SachanEMNLP 2023 · 11 citations
- Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in TransformersAndrew Nam, Henry Conklin, Yukang Yang, Tom Griffiths et al.NeurIPS 2025 · 21 citations
- Causal Interpretation of Self-Attention in Pre-Trained TransformersRaanan Y. Rohekar, Yaniv Gurwicz, Shami NisimovNeurIPS 2023 · 62 citations
