Investigating Gender Bias in Language Models Using Causal Mediation Analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, Stuart M. Shieber
摘要
Many interpretation methods for neural models in natural language processing investigate how information is encoded inside hidden representations. However, these methods can only measure whether the information exists, not whether it is actually used by the model. We propose a methodology grounded in the theory of causal mediation analysis for interpreting which parts of a model are causally implicated in its behavior. The approach enables us to analyze the mechanisms that facilitate the flow of information from input to output through various model components, known as mediators. As a case study, we apply this methodology to analyzing gender bias in pre-trained Transformer language models. We study the role of individual neurons and attention heads in mediating gender bias across three datasets designed to gauge a model's sensitivity to gender bias. Our mediation analysis reveals that gender bias effects are concentrated in specific components of the model that may exhibit highly specialized behavior.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper199
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 被引用 651 次
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong 等ICLR 2021 · 被引用 357 次
它引用的顶会 Paper4
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- A Causal Inference Method for Reducing Gender Bias in Word Embedding RelationsZekun Yang, Juan FengAAAI 2020 · 被引用 40 次
- Toward Gender-Inclusive Coreference ResolutionYang Trista Cao, Hal Daumé IIIACL 2020 · 被引用 20 次
相关 Paper
- Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation PerspectiveZhaotian Weng, Zijun Gao, Jerone Theodore Alexander Andrews, Jieyu ZhaoEMNLP 2024 · 被引用 3 次
- Debiasing Algorithm through Model AdaptationTomasz Limisiewicz, David Marecek, Tomás MusilICLR 2024 · 被引用 24 次
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisAlessandro Stolfo, Yonatan Belinkov, Mrinmaya SachanEMNLP 2023 · 被引用 11 次
- Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in TransformersAndrew Nam, Henry Conklin, Yukang Yang, Tom Griffiths 等NeurIPS 2025 · 被引用 21 次
- Causal Interpretation of Self-Attention in Pre-Trained TransformersRaanan Y. Rohekar, Yaniv Gurwicz, Shami NisimovNeurIPS 2023 · 被引用 62 次
