Linear Adversarial Concept Erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan Cotterell
摘要
Modern neural models trained on textual data rely on pre-trained representations that emerge without direct supervision. As these representations are increasingly being used in real-world applications, the inability to control their content becomes an increasingly important problem. We formulate the problem of identifying and erasing a linear subspace that corresponds to a given concept, in order to prevent linear predictors from recovering the concept. We model this problem as a constrained, linear maximin game, and show that existing solutions are generally not optimal for this task. We derive a closed-form solution for certain objectives, and propose a convex relaxation, , that works well for others. When evaluated in the context of binary gender removal, the method recovers a low-dimensional subspace whose removal mitigates bias by intrinsic and extrinsic evaluation. We show that the method is highly expressive, effectively mitigating bias in deep nonlinear classifiers while maintaining tractability and interpretability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell 等NeurIPS 2023 · 被引用 305 次
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger 等NeurIPS 2024 · 被引用 233 次
- In-Context Unlearning: Language Models as Few-Shot UnlearnersMartin Pawelczyk, Seth Neel, Himabindu LakkarajuICML 2024 · 被引用 217 次
- Probing for the Usage of Grammatical NumberKarim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau 等ACL 2022 · 被引用 72 次
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 被引用 46 次
它引用的顶会 Paper6
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell 等NeurIPS 2023 · 被引用 305 次
- On the Global Optima of Kernelized Adversarial Representation LearningBashir Sadeghi, Runyi Yu, Vishnu BoddetiICCV 2019 · 被引用 34 次
- OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarEMNLP 2021 · 被引用 31 次
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton 等ACL 2020 · 被引用 25 次
- The complexity of constrained min-max optimizationConstantinos Daskalakis, Stratis Skoulakis, Manolis ZampetakisSTOC 2021 · 被引用 18 次
相关 Paper
- Adversarial Concept Erasure in Kernel SpaceShauli Ravfogel, Francisco Vargas, Yoav Goldberg, Ryan CotterellEMNLP 2022 · 被引用 11 次
- Linear Guardedness and its ImplicationsShauli Ravfogel, Yoav Goldberg, Ryan CotterellACL 2023 · 被引用 2 次
- Exploring the Linear Subspace Hypothesis in Gender Bias MitigationFrancisco Vargas, Ryan CotterellEMNLP 2020 · 被引用 2 次
- Prototype-Guided Concept Erasure in Diffusion ModelsYuze Cai, Jiahao Lu, Hongxiang Shi, Yichao Zhou 等CVPR 2026 · 被引用 3 次
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 被引用 4 次
