Linear Guardedness and its Implications
Shauli Ravfogel, Yoav Goldberg, Ryan Cotterell
摘要
Methods for erasing human-interpretable concepts from neural representations that assume linearity have been found to be tractable and useful.However, the impact of this removal on the behavior of downstream classifiers trained on the modified representations is not fully understood.In this work, we formally define the notion of linear guardedness as the inability of an adversary to predict the concept directly from the representation, and study its implications.We show that, in the binary case, under certain assumptions, a downstream log-linear model cannot recover the erased concept.However, we constructively demonstrate that a multiclass log-linear model can be constructed that indirectly recovers the concept in some cases, pointing to the inherent limitations of linear guardedness as a downstream bias mitigation technique.These findings shed light on the theoretical limitations of linear erasure methods and highlight the need for further research on the connections between intrinsic and extrinsic bias in neural models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell 等NeurIPS 2023 · 被引用 305 次
- Representation Surgery: Theory and Practice of Affine SteeringShashwat Singh, Shauli Ravfogel, Jonathan Herzig, Roee Aharoni 等ICML 2024 · 被引用 36 次
- The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language ModelsAviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan 等EMNLP 2023 · 被引用 11 次
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 被引用 4 次
- Language Concept Erasure for Language-invariant Dense RetrievalZhiqi Huang, Puxuan Yu, Shauli Ravfogel, James AllanEMNLP 2024 · 被引用 1 次
它引用的顶会 Paper7
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell 等NeurIPS 2023 · 被引用 305 次
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart 等ICLR 2020 · 被引用 211 次
- Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language ModelsRyan Steed, Swetasudha Panda, Ari Kobren, Michael L. WickACL 2022 · 被引用 52 次
- OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarEMNLP 2021 · 被引用 31 次
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton 等ACL 2020 · 被引用 25 次
相关 Paper
- Linear Adversarial Concept ErasureShauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan CotterellICML 2022 · 被引用 89 次
- Adversarial Concept Erasure in Kernel SpaceShauli Ravfogel, Francisco Vargas, Yoav Goldberg, Ryan CotterellEMNLP 2022 · 被引用 11 次
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 被引用 46 次
- Obliviator Reveals the Cost of Nonlinear Guardedness in Concept ErasureRamin Akbari, Milad Afshari, Vishnu BoddetiNeurIPS 2025 · 被引用 2 次
- On the Origins of Linear Representations in Large Language ModelsYibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam 等ICML 2024 · 被引用 68 次
