Underestimated Privacy Risks for Minority Populations in Large Language Model Unlearning
Rongzhe Wei, Mufei Li, Mohsen Ghassemi, Eleonora Kreacic, Yifan Li, Xiang Yue, Bo Li, Vamsi K. Potluru, Pan Li, Eli Chien
摘要
Large Language Models (LLMs) embed sensitive, human-generated data, prompting the need for unlearning methods. Although certified unlearning offers strong privacy guarantees, its restrictive assumptions make it unsuitable for LLMs, giving rise to various heuristic approaches typically assessed through empirical evaluations. These standard evaluations randomly select data for removal, apply unlearning techniques, and use membership inference attacks (MIAs) to compare unlearned models against models retrained without the removed data. However, to ensure robust privacy protections for every data point, it is essential to account for scenarios in which certain data subsets face elevated risks. Prior research suggests that outliers, particularly including data tied to minority groups, often exhibit higher memorization propensity which indicates they may be more difficult to unlearn. Building on these insights, we introduce a complementary, minority-aware evaluation framework to highlight blind spots in existing frameworks. We substantiate our findings with carefully designed experiments, using canaries with personally identifiable information (PII) to represent these minority subsets and demonstrate that they suffer at least 20% higher privacy leakage across various unlearning methods, MIAs, datasets, and LLM scales. Our proposed minorityaware evaluation framework 1 marks an essential step toward more equitable and comprehensive assessments of LLM unlearning efficacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence AwarenessRongzhe Wei, Peizhi Niu, Hans Hao-Hsun Hsu, Ruihan Wu 等NeurIPS 2025 · 被引用 20 次
- ReLearn: Unlearning via Learning for Large Language ModelsHaoming Xu, Ningyuan Zhao, Liming Yang, Sendong Zhao 等ACL 2025 · 被引用 18 次
- MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMsYupu Gu, Rongzhe Wei, Andy Zhu, Pan LiICLR 2026 · 被引用 4 次
它引用的顶会 Paper38
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
相关 Paper
- Information-Theoretic Membership Inference for Granular Quantification of MemorizationJiashu Tao, Reza ShokriICLR 2026
- Identifying Unlearned Data in LLMs via Membership Inference AttacksAdvit Deepak, Megan Mou, Jing Huang, Diyi YangEMNLP 2025
- A Reliable Cryptographic Framework for Empirical Machine Unlearning EvaluationYiwen Tu, Pingbang Hu, Jiaqi MaNeurIPS 2025 · 被引用 6 次
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language ModelsXiaoyu Xu, Minxin Du, Qingqing Ye, Haibo HuEMNLP 2025 · 被引用 1 次
- Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization FrameworkXiaoyu Luo, Yiyi Chen, Qiongxiu Li, Johannes BjervaACL 2026 · 被引用 1 次
