Linear Guardedness and its Implications
Shauli Ravfogel, Yoav Goldberg, Ryan Cotterell
Abstract
Methods for erasing human-interpretable concepts from neural representations that assume linearity have been found to be tractable and useful.However, the impact of this removal on the behavior of downstream classifiers trained on the modified representations is not fully understood.In this work, we formally define the notion of linear guardedness as the inability of an adversary to predict the concept directly from the representation, and study its implications.We show that, in the binary case, under certain assumptions, a downstream log-linear model cannot recover the erased concept.However, we constructively demonstrate that a multiclass log-linear model can be constructed that indirectly recovers the concept in some cases, pointing to the inherent limitations of linear guardedness as a downstream bias mitigation technique.These findings shed light on the theoretical limitations of linear erasure methods and highlight the need for further research on the connections between intrinsic and extrinsic bias in neural models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2886e8fd-435a-48df-abba-043faec566dfCited by top-tier papers8
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell et al.NeurIPS 2023 · 305 citations
- Representation Surgery: Theory and Practice of Affine SteeringShashwat Singh, Shauli Ravfogel, Jonathan Herzig, Roee Aharoni et al.ICML 2024 · 36 citations
- The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language ModelsAviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan et al.EMNLP 2023 · 11 citations
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 4 citations
- Language Concept Erasure for Language-invariant Dense RetrievalZhiqi Huang, Puxuan Yu, Shauli Ravfogel, James AllanEMNLP 2024 · 1 citation
Builds on7
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell et al.NeurIPS 2023 · 305 citations
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart et al.ICLR 2020 · 211 citations
- Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language ModelsRyan Steed, Swetasudha Panda, Ari Kobren, Michael L. WickACL 2022 · 52 citations
- OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarEMNLP 2021 · 31 citations
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton et al.ACL 2020 · 25 citations
Related papers
- Linear Adversarial Concept ErasureShauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan CotterellICML 2022 · 89 citations
- Adversarial Concept Erasure in Kernel SpaceShauli Ravfogel, Francisco Vargas, Yoav Goldberg, Ryan CotterellEMNLP 2022 · 11 citations
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 46 citations
- Obliviator Reveals the Cost of Nonlinear Guardedness in Concept ErasureRamin Akbari, Milad Afshari, Vishnu BoddetiNeurIPS 2025 · 2 citations
- On the Origins of Linear Representations in Large Language ModelsYibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam et al.ICML 2024 · 68 citations
