SEPS: A Separability Measure for Robust Unlearning in LLMs
Wonje Jeung, Sangyeon Yoon, Albert No
摘要
Machine unlearning aims to selectively remove targeted knowledge from Large Language Models (LLMs), ensuring they forget specified content while retaining essential information.Existing unlearning metrics assess whether a model correctly answers retain queries and rejects forget queries, but they fail to capture real-world scenarios where forget queries rarely appear in isolation.In fact, forget and retain queries often coexist within the same prompt, making mixed-query evaluation crucial.We introduce SEPS, an evaluation framework that explicitly measures a model's ability to both forget and retain information within a single prompt.Through extensive experiments across three benchmarks, we identify two key failure modes in existing unlearning methods: (1) untargeted unlearning indiscriminately erases both forget and retain content once a forget query appears, and (2) targeted unlearning overfits to single-query scenarios, leading to catastrophic failures when handling multiple queries.To address these issues, we propose Mixed Prompt (MP) unlearning, a strategy that integrates both forget and retain queries into a unified training objective.Our approach significantly improves unlearning effectiveness, demonstrating robustness even in complex settings with up to eight mixed forget and retain queries in a single prompt.We release code at https://github.com/AI-ISL/SEPS.Table 9: Prompt for generating WMDP retain QA pairs for biology, chemistry, and cybersecurity datasets.Generate 100 biology MCQs in this format: "question": "", "answer": "", "choices": ["","","",""] Requirements: * Middle school level * Focus on core concepts (cells, ecosystems, human body) * Safe content (no harmful techniques) * Include real-world examples * Use simple, clear language * answer must be 0-3 Generate 100 chemistry MCQs in this format: "question": "", "answer": "", "choices": ["","","",""] Requirements: * Middle school level * Focus on basic concepts (atoms, reactions, states of matter) * No dangerous chemical combinations * Clear, simple language * answer must be 0-3 Generate 100 cybersecurity MCQs in this format: "question": "", "answer": "", "choices": ["","","",""] Requirements: * Middle school level * Safe content (no harmful techniques) * Focus on practical digital safety (passwords, privacy, online behavior) * Include real-world examples * Use simple, clear language
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning FailuresSangyeon Yoon, Hyesoo Hong, Wonje Jeung, Albert NoICLR 2026 · 被引用 3 次
- R-TOFU: Unlearning in Large Reasoning ModelsSangyeon Yoon, Wonje Jeung, Albert NoEMNLP 2025 · 被引用 1 次
- DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language ModelsXuyang Zhong, Qizhang Li, Yiwen Guo, Chen LiuICML 2026
- De-attribute to Forget for LLM UnlearningXinyang Lu, Jiabao Pan, Rachael Hwee Ling Sim, See-Kiong Ng 等ICML 2026
它引用的顶会 Paper23
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song 等S&P 2022 · 被引用 1,049 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 被引用 516 次
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue 等ICML 2024 · 被引用 390 次
相关 Paper
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language ModelsYuefeng Peng, Parnian Afshar, Megan Ganji, Thomas Butler 等ICML 2026 · 被引用 1 次
- Towards Effective Evaluations and Comparisons for LLM Unlearning MethodsQizhou Wang, Bo Han, Puning Yang, Jianing Zhu 等ICLR 2025
- A Closer Look at Machine Unlearning for Large Language ModelsXiaojian Yuan, Tianyu Pang, Chao Du, Kejiang Chen 等ICLR 2025
- MUSE: Machine Unlearning Six-Way Evaluation for Language ModelsWeijia Shi, Jaechan Lee, Yangsibo Huang, Sadhika Malladi 等ICLR 2025
- OFMU: Optimization-Driven Framework for Machine UnlearningSadia Asif, Mohammad Mohammadi AmiriICLR 2026 · 被引用 4 次
