Underestimated Privacy Risks for Minority Populations in Large Language Model Unlearning
Rongzhe Wei, Mufei Li, Mohsen Ghassemi, Eleonora Kreacic, Yifan Li, Xiang Yue, Bo Li, Vamsi K. Potluru, Pan Li, Eli Chien
Abstract
Large Language Models (LLMs) embed sensitive, human-generated data, prompting the need for unlearning methods. Although certified unlearning offers strong privacy guarantees, its restrictive assumptions make it unsuitable for LLMs, giving rise to various heuristic approaches typically assessed through empirical evaluations. These standard evaluations randomly select data for removal, apply unlearning techniques, and use membership inference attacks (MIAs) to compare unlearned models against models retrained without the removed data. However, to ensure robust privacy protections for every data point, it is essential to account for scenarios in which certain data subsets face elevated risks. Prior research suggests that outliers, particularly including data tied to minority groups, often exhibit higher memorization propensity which indicates they may be more difficult to unlearn. Building on these insights, we introduce a complementary, minority-aware evaluation framework to highlight blind spots in existing frameworks. We substantiate our findings with carefully designed experiments, using canaries with personally identifiable information (PII) to represent these minority subsets and demonstrate that they suffer at least 20% higher privacy leakage across various unlearning methods, MIAs, datasets, and LLM scales. Our proposed minorityaware evaluation framework 1 marks an essential step toward more equitable and comprehensive assessments of LLM unlearning efficacy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 965e340d-d0a9-4e36-b4d0-bce4d3ec195fCited by top-tier papers3
- Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence AwarenessRongzhe Wei, Peizhi Niu, Hans Hao-Hsun Hsu, Ruihan Wu et al.NeurIPS 2025 · 20 citations
- ReLearn: Unlearning via Learning for Large Language ModelsHaoming Xu, Ningyuan Zhao, Liming Yang, Sendong Zhao et al.ACL 2025 · 18 citations
- MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMsYupu Gu, Rongzhe Wei, Andy Zhu, Pan LiICLR 2026 · 4 citations
Builds on38
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
Related papers
- Information-Theoretic Membership Inference for Granular Quantification of MemorizationJiashu Tao, Reza ShokriICLR 2026
- Identifying Unlearned Data in LLMs via Membership Inference AttacksAdvit Deepak, Megan Mou, Jing Huang, Diyi YangEMNLP 2025
- A Reliable Cryptographic Framework for Empirical Machine Unlearning EvaluationYiwen Tu, Pingbang Hu, Jiaqi MaNeurIPS 2025 · 6 citations
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language ModelsXiaoyu Xu, Minxin Du, Qingqing Ye, Haibo HuEMNLP 2025 · 1 citation
- Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization FrameworkXiaoyu Luo, Yiyi Chen, Qiongxiu Li, Johannes BjervaACL 2026 · 1 citation
