Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language Models
Elena Sofia Ruzzetti, Giancarlo A. Xompero, Davide Venditti, Fabio Massimo Zanzotto
摘要
Large Language Models (LLMs) memorize, and thus, among huge amounts of uncontrolled data, may memorize Personally Identifiable Information (PII), which should not be stored and, consequently, not leaked. In this paper, we introduce Private Memorization Editing (PME), an approach for preventing private data leakage that turns an apparent limitation, that is, the LLMs' memorization ability, into a powerful privacy defense strategy. While attacks against LLMs have been performed exploiting previous knowledge regarding their training data, our approach aims to exploit the same kind of knowledge in order to make a model more robust. We detect a memorized PII and then mitigate the memorization of PII by editing a model knowledge of its training data. We verify that our procedure does not affect the underlying language model while making it more robust against privacy Training Data Extraction attacks. We demonstrate that PME can effectively reduce the number of leaked PII in a number of configurations, in some cases even reducing the accuracy of the privacy attacks to zero.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement LearningZheng Hui, Yijiang River Dong, Sanhanat Sivapiromrat, Ehsan Shareghi 等ACL 2026 · 被引用 4 次
- Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization FrameworkXiaoyu Luo, Yiyi Chen, Qiongxiu Li, Johannes BjervaACL 2026 · 被引用 1 次
它引用的顶会 Paper15
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 被引用 365 次
相关 Paper
- Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model UtilityMartin Kuo, Jingyang Zhang, Jianyi Zhang, Minxue Tang 等ICLR 2025
- Teach LLMs to Phish: Stealing Private Information from Language ModelsAshwinee Panda, Christopher A. Choquette-Choo, Zhengming Zhang, Yaoqing Yang 等ICLR 2024 · 被引用 41 次
- Rethinking the Role of Verbatim Memorization in LLM PrivacyTom Sander, Bargav Jayaraman, Mark Ibrahim, Kamalika Chaudhuri 等NeurIPS 2025 · 被引用 5 次
- Private Investigator: Extracting Personally Identifiable Information from Large Language Models Using Optimized PromptsSeongho Keum, Dongwon Shin, Leo Marchyok, Sanghyun Hong 等USENIX Security 2025
- Quantifying and Analyzing Entity-Level Memorization in Large Language ModelsZhenhong Zhou, Jiuyang Xiang, Chaomeng Chen, Sen SuAAAI 2024 · 被引用 23 次
