Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks
Vaidehi Patil, Peter Hase, Mohit Bansal
摘要
Pretrained language models sometimes possess knowledge that we do not wish them to, including memorized personal information and knowledge that could be used to harm people. They can also output toxic or harmful text. To mitigate these safety and informational issues, we propose an attack-and-defense framework for studying the task of deleting sensitive information directly from model weights. We study direct edits to model weights because (1) this approach should guarantee that particular deleted information is never extracted by future prompt attacks, and (2) it should protect against whitebox attacks, which is necessary for making claims about safety/privacy in a setting where publicly available model weights could be used to elicit sensitive information. Our threat model assumes that an attack succeeds if the answer to a sensitive question is located among a set of B generated candidates, based on scenarios where the information would be insecure if the answer is among B candidates. Experimentally, we show that even state-of-the-art model editing methods such as ROME struggle to truly delete factual information from models like GPT-J, as our whitebox and blackbox attacks can recover "deleted" information from an edited model 38% of the time. These attacks leverage two key observations: (1) that traces of deleted information can be found in intermediate model hidden states, and (2) that applying an editing method for one question may not delete information across rephrased versions of the question. Finally, we provide new defense methods that protect against some extraction attacks, but we do not find a single universally effective defense method. Our results suggest that truly deleting sensitive information is a tractable but difficult problem, since even relatively low attack success rates have potentially severe implications for the deployment of language models in a world where individuals enjoy ownership of their personal data, a right to privacy, and safety from harmful model outputs. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper52
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and GenerationChongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong 等ICLR 2024 · 被引用 351 次
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM UnlearningChongyu Fan, Jiancheng Liu, Licong Lin, Jinghan Jia 等NeurIPS 2025 · 被引用 182 次
- Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding SpaceLeo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel 等NeurIPS 2024 · 被引用 113 次
- Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language ModelsLingzhi Wang, Xingshan Zeng, Jinsong Guo, Kam-Fai Wong 等AAAI 2025 · 被引用 43 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
相关 Paper
- Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language ModelsElena Sofia Ruzzetti, Giancarlo A. Xompero, Davide Venditti, Fabio Massimo ZanzottoACL 2025 · 被引用 9 次
- Concept-ROT: Poisoning Concepts in Large Language Models with Model EditingKeltin Grimes, Marco Christiani, David Shriver, Marissa Catherine ConnorICLR 2025
- Large Scale Knowledge WashingYu Wang, Ruihan Wu, Zexue He, Xiusi Chen 等ICLR 2025
- MoPe: Model Perturbation based Privacy Attacks on Language ModelsMarvin Li, Jason Wang, Jeffrey G. Wang, Seth NeelEMNLP 2023 · 被引用 8 次
- Tracing and Reversing Edits in LLMsPaul Youssef, Zhixue Zhao, Christin Seifert, Jörg SchlöttererICLR 2026 · 被引用 7 次
