LEACE: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, Stella Biderman
摘要
Concept erasure aims to remove specified features from an embedding. It can improve fairness (e.g. preventing a classifier from using gender or race) and interpretability (e.g. removing a concept to observe changes in model behavior). We introduce LEAst-squares Concept Erasure (LEACE), a closed-form method which provably prevents all linear classifiers from detecting a concept while changing the embedding as little as possible, as measured by a broad class of norms. We apply LEACE to large language models with a novel procedure called concept scrubbing, which erases target concept information from every layer in the network. We demonstrate our method on two tasks: measuring the reliance of language models on part-of-speech information, and reducing gender bias in BERT embeddings. Our code is available at https://github.com/EleutherAI/concept-erasure .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper61
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue 等ICML 2024 · 被引用 390 次
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelMichael Hanna, Ollie Liu, Alexandre VariengienNeurIPS 2023 · 被引用 251 次
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger 等NeurIPS 2024 · 被引用 233 次
- Large Language Model Unlearning via Embedding-Corrupted PromptsChris Yuhao Liu, Yaxuan Wang, Jeffrey Flanigan, Yang LiuNeurIPS 2024 · 被引用 138 次
它引用的顶会 Paper13
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart 等ICLR 2020 · 被引用 211 次
- Interpretability at Scale: Identifying Causal Mechanisms in AlpacaZhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts 等NeurIPS 2023 · 被引用 146 次
- Inducing Causal Structure for Interpretable Neural NetworksAtticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner 等ICML 2022 · 被引用 104 次
- Linear Adversarial Concept ErasureShauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan CotterellICML 2022 · 被引用 89 次
- Causal Conceptions of Fairness and their ConsequencesHamed Nilforoshan, Johann D. Gaebler, Ravi Shroff, Sharad GoelICML 2022 · 被引用 52 次
相关 Paper
- Precise In-Parameter Concept Erasure in Large Language ModelsYoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez 等EMNLP 2025 · 被引用 10 次
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 被引用 4 次
- Adversarial Concept Erasure in Kernel SpaceShauli Ravfogel, Francisco Vargas, Yoav Goldberg, Ryan CotterellEMNLP 2022 · 被引用 11 次
- Prototype-Guided Concept Erasure in Diffusion ModelsYuze Cai, Jiahao Lu, Hongxiang Shi, Yichao Zhou 等CVPR 2026 · 被引用 3 次
- Circumventing Concept Erasure Methods For Text-To-Image Generative ModelsMinh Pham, Kelly O. Marshall, Niv Cohen, Govind Mittal 等ICLR 2024 · 被引用 82 次
