SEPS: A Separability Measure for Robust Unlearning in LLMs
Wonje Jeung, Sangyeon Yoon, Albert No
Abstract
Machine unlearning aims to selectively remove targeted knowledge from Large Language Models (LLMs), ensuring they forget specified content while retaining essential information.Existing unlearning metrics assess whether a model correctly answers retain queries and rejects forget queries, but they fail to capture real-world scenarios where forget queries rarely appear in isolation.In fact, forget and retain queries often coexist within the same prompt, making mixed-query evaluation crucial.We introduce SEPS, an evaluation framework that explicitly measures a model's ability to both forget and retain information within a single prompt.Through extensive experiments across three benchmarks, we identify two key failure modes in existing unlearning methods: (1) untargeted unlearning indiscriminately erases both forget and retain content once a forget query appears, and (2) targeted unlearning overfits to single-query scenarios, leading to catastrophic failures when handling multiple queries.To address these issues, we propose Mixed Prompt (MP) unlearning, a strategy that integrates both forget and retain queries into a unified training objective.Our approach significantly improves unlearning effectiveness, demonstrating robustness even in complex settings with up to eight mixed forget and retain queries in a single prompt.We release code at https://github.com/AI-ISL/SEPS.Table 9: Prompt for generating WMDP retain QA pairs for biology, chemistry, and cybersecurity datasets.Generate 100 biology MCQs in this format: "question": "", "answer": "", "choices": ["","","",""] Requirements: * Middle school level * Focus on core concepts (cells, ecosystems, human body) * Safe content (no harmful techniques) * Include real-world examples * Use simple, clear language * answer must be 0-3 Generate 100 chemistry MCQs in this format: "question": "", "answer": "", "choices": ["","","",""] Requirements: * Middle school level * Focus on basic concepts (atoms, reactions, states of matter) * No dangerous chemical combinations * Clear, simple language * answer must be 0-3 Generate 100 cybersecurity MCQs in this format: "question": "", "answer": "", "choices": ["","","",""] Requirements: * Middle school level * Safe content (no harmful techniques) * Focus on practical digital safety (passwords, privacy, online behavior) * Include real-world examples * Use simple, clear language
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f77823ad-688f-4e94-9b85-8fcee83857dcCited by top-tier papers4
- Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning FailuresSangyeon Yoon, Hyesoo Hong, Wonje Jeung, Albert NoICLR 2026 · 3 citations
- R-TOFU: Unlearning in Large Reasoning ModelsSangyeon Yoon, Wonje Jeung, Albert NoEMNLP 2025 · 1 citation
- DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language ModelsXuyang Zhong, Qizhang Li, Yiwen Guo, Chen LiuICML 2026
- De-attribute to Forget for LLM UnlearningXinyang Lu, Jiabao Pan, Rachael Hwee Ling Sim, See-Kiong Ng et al.ICML 2026
Builds on23
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
Related papers
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language ModelsYuefeng Peng, Parnian Afshar, Megan Ganji, Thomas Butler et al.ICML 2026 · 1 citation
- Towards Effective Evaluations and Comparisons for LLM Unlearning MethodsQizhou Wang, Bo Han, Puning Yang, Jianing Zhu et al.ICLR 2025
- A Closer Look at Machine Unlearning for Large Language ModelsXiaojian Yuan, Tianyu Pang, Chao Du, Kejiang Chen et al.ICLR 2025
- MUSE: Machine Unlearning Six-Way Evaluation for Language ModelsWeijia Shi, Jaechan Lee, Yangsibo Huang, Sadhika Malladi et al.ICLR 2025
- OFMU: Optimization-Driven Framework for Machine UnlearningSadia Asif, Mohammad Mohammadi AmiriICLR 2026 · 4 citations
