Reinforcement Unlearning
Dayong Ye, Tianqing Zhu, Congcong Zhu, Derui Wang, Kun Gao, Zewei Shi, Sheng Shen, Wanlei Zhou, Minhui Xue
摘要
The widespread deployment of Large Language Models (LLMs) trained on massive, uncurated corpora has raised growing concerns about the inclusion of sensitive, copyrighted, or illegal content. This has led to increasing interest in LLM unlearning: the task of selectively removing specific information from a model without retraining from scratch or degrading overall utility. However, existing methods often rely on large-scale forget and retain datasets, and suffer from unnatural responses, poor generalization, or catastrophic utility loss. In this work, we propose Reinforcement UnLearning (RULE), an efficient framework that formulates unlearning as a refusal boundary optimization problem. RULE is trained with a small portion of the forget set and synthesized boundary queries, using a verifiable reward function that encourages safe refusal on forget--related queries while preserving helpful responses on permissible inputs. We provide both theoretical and empirical evidence demonstrating the effectiveness of RULE in achieving targeted unlearning without compromising model utility. Experimental results show that, with only forget set and synthesized boundary data, RULE outperforms existing baselines by up to forget quality and naturalness response while maintaining general utility, achieving forget--retain Pareto optimality. Remarkably, we further observe that RULE improves the naturalness of model outputs, enhances training efficiency, and exhibits strong generalization ability, generalizing refusal behavior to semantically related but unseen queries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement LearningChen Gong, Zheng Liu, Kecen Li, Tianhao WangNDSS 2026 · 被引用 3 次
- Unlearning’s Blind Spots: Over‑Unlearning and Prototypical Relearning AttackSeungBum Ha, Saerom Park, Sung Whan YoonICML 2026 · 被引用 2 次
- TrajDeleter: Enabling Trajectory Forgetting in Offline Reinforcement Learning AgentsChen Gong, Kecen Li, Jin Yao, Tianhao WangNDSS 2025
它引用的顶会 Paper19
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
- Towards Unbounded Machine UnlearningMeghdad Kurmanji, Peter Triantafillou, Jamie Hayes, Eleni TriantafillouNeurIPS 2023 · 被引用 363 次
- Adaptive Machine UnlearningVarun Gupta, Christopher Jung, Seth Neel, Aaron Roth 等NeurIPS 2021 · 被引用 262 次
相关 Paper
- RULE: Reinforcement UnLEarning Achieves Forget-retain Pareto OptimalityChenlong Zhang, Zhuoran Jin, Hongbang Yuan, Jiaheng Wei 等NeurIPS 2025 · 被引用 15 次
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language ModelsYuefeng Peng, Parnian Afshar, Megan Ganji, Thomas Butler 等ICML 2026 · 被引用 1 次
- CAP: Controllable Alignment Prompting for Unlearning in LLMsZhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu 等ACL 2026
- De-attribute to Forget for LLM UnlearningXinyang Lu, Jiabao Pan, Rachael Hwee Ling Sim, See-Kiong Ng 等ICML 2026
- ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language ModelsJiahui Guang, Haiyan Wang, Yingjie Zhu, Cuiyun Gao 等ICML 2026
