Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, Yoav Goldberg
Abstract
The ability to control for the kinds of information encoded in neural representation has a variety of use cases, especially in light of the challenge of interpreting these models. We present Iterative Null-space Projection (INLP), a novel method for removing information from neural representations. Our method is based on repeated training of linear classifiers that predict a certain property we aim to remove, followed by projection of the representations on their null-space. By doing so, the classifiers become oblivious to that target property, making it hard to linearly separate the data according to it. While applicable for multiple uses, we evaluate our method on bias and fairness use-cases, and show that our method is able to mitigate bias in word embeddings, as well as to increase fairness in a setting of multi-class classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70106a2e-c5f4-4b9b-97ac-e51a98ed6e59Cited by top-tier papers122
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 495 citations
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell et al.NeurIPS 2023 · 305 citations
- Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction AttacksVaidehi Patil, Peter Hase, Mohit BansalICLR 2024 · 167 citations
- An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language ModelsNicholas Meade, Elinor Poole-Dayan, Siva ReddyACL 2022 · 160 citations
Related papers
- Better Hit the Nail on the Head than Beat around the Bush: Removing Protected Attributes with a Single ProjectionPantea Haghighatkhah, Antske Fokkens, Pia Sommerauer, Bettina Speckmann et al.EMNLP 2022 · 2 citations
- Linear Adversarial Concept ErasureShauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan CotterellICML 2022 · 89 citations
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 4 citations
- Interpretable Debiasing of Vectorized Language Representations with Iterative OrthogonalizationPrince Osei Aboagye, Yan Zheng, Jack Shunn, Chin-Chia Michael Yeh et al.ICLR 2023
- Removing Spurious Concepts from Neural Network Representations via Joint Subspace EstimationFloris Holstege, Bram Wouters, Noud P. A. van Giersbergen, Cees G. H. DiksICML 2024 · 3 citations
