Textual Unlearning Gives a False Sense of Unlearning
Jiacheng Du, Zhibo Wang, Jie Zhang, Xiaoyi Pang, Jiahui Hu, Kui Ren
摘要
Language Models (LMs) are prone to "memorizing" training data, including substantial sensitive user information. To mitigate privacy risks and safeguard the right to be forgotten, machine unlearning has emerged as a promising approach for enabling LMs to efficiently "forget" specific texts. However, despite the good intentions, is textual unlearning really as effective and reliable as expected? To address the concern, we first propose Unlearning Likelihood Ratio Attack+ (U-LiRA+), a rigorous textual unlearning auditing method, and find that unlearned texts can still be detected with very high confidence after unlearning. Further, we conduct an in-depth investigation on the privacy risks of textual unlearning mechanisms in deployment and present the Textual Unlearning Leakage Attack (TULA), along with its variants in both black-and white-box scenarios. We show that textual unlearning mechanisms could instead reveal more about the unlearned texts, exposing them to significant membership inference and data reconstruction risks. Our findings highlight that existing textual unlearning actually gives a false sense of unlearning, underscoring the need for more robust and secure unlearning mechanisms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 被引用 7 次
- Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials SciencePeter A. Jansen, Samiah Hassan, Ruoyao WangEMNLP 2025 · 被引用 4 次
- Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language ModelsHuazheng Wang, Yongcheng Jing, Haifeng Sun, Yingjie Wang 等ACL 2026 · 被引用 3 次
- Towards Effective Evaluations and Comparisons for LLM Unlearning MethodsQizhou Wang, Bo Han, Puning Yang, Jianing Zhu 等ICLR 2025
- Multilingual Unlearning in LLMs: Transfer, Dynamics, and ReversibilityChaoyi Xiang, Olga Ohrimenko, Benjamin Rubinstein, Lea FrermannICML 2026
它引用的顶会 Paper15
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
相关 Paper
- When Machine Unlearning Jeopardizes PrivacyMin Chen, Zhikun Zhang, Tianhao Wang, Michael Backes 等CCS 2021 · 被引用 146 次
- Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine UnlearningHongsheng Hu, Shuo Wang, Tian Dong, Minhui XueS&P 2024 · 被引用 62 次
- Identifying Unlearned Data in LLMs via Membership Inference AttacksAdvit Deepak, Megan Mou, Jing Huang, Diyi YangEMNLP 2025
- Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for PrivacyYaxin Xiao, Qingqing Ye, Li Hu, Huadi Zheng 等ICCV 2025 · 被引用 6 次
- TAPE: Tailored Posterior Difference for Auditing of Machine UnlearningWeiqi Wang, Zhiyi Tian, An Liu, Shui YuWWW 2025 · 被引用 9 次
