On Effects of Steering Latent Representation for Large Language Model Unlearning
Huu-Tien Dang, Tin Pham, Hoang Thanh-Tung, Naoya Inoue
摘要
Representation Misdirection for Unlearning (RMU), which steers model representation in the intermediate layer to a target random representation, is an effective method for large language model (LLM) unlearning. Despite its high performance, the underlying cause and explanation remain underexplored. In this paper, we theoretically demonstrate that steering forget representations in the intermediate layer reduces token confidence, causing LLMs to generate wrong or nonsense responses. We investigate how the coefficient influences the alignment of forget-sample representations with the random direction and hint at the optimal coefficient values for effective unlearning across different network layers. We show that RMU unlearned models are robust against adversarial jailbreak attacks. Furthermore, our empirical analysis shows that RMU is less effective when applied to the middle and later layers in LLMs. To resolve this drawback, we propose Adaptive RMU-a simple yet effective alternative method that makes unlearning effective with most layers. Extensive experiments demonstrate that Adaptive RMU significantly improves the unlearning performance compared to prior art while incurring no additional computational cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Keeping an Eye on LLM Unlearning: The Hidden Risk and RemedyJie Ren, Zhenwei Dai, Xianfeng Tang, Yue Xing 等NeurIPS 2025 · 被引用 11 次
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 被引用 7 次
- LLM Unlearning Should Be Form-IndependentXiaotian Ye, Mengqi Zhang, Shu WuS&P 2026 · 被引用 3 次
- Randomized Antipodal Search Done Right for Data Pareto Improvement of LLM UnlearningZiwen Liu, Huawei Lin, Yide Ran, Denghui Zhang 等ICLR 2026 · 被引用 2 次
- Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language ModelsKunhao Li, Wenhao Li, Di Wu, Lei Yang 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper36
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 被引用 789 次
- Scaling Out-of-Distribution Detection for Real-World SettingsDan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou 等ICML 2022 · 被引用 653 次
相关 Paper
- Model Unlearning via Sparse Autoencoder Subspace Guided ProjectionsXu Wang, Zihao Li, Benyou Wang, Yan Hu 等EMNLP 2025 · 被引用 9 次
- JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path RectificationXi Wang, Songlei Jian, Shasha Li, Xiaopeng Li 等ACL 2026 · 被引用 2 次
- Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning SkillsChangsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia 等EMNLP 2025
- LLM Unlearning via Neural Activation RedirectionWilliam F. Shen, Xinchi Qiu, Meghdad Kurmanji, Alex Iacob 等NeurIPS 2025 · 被引用 1 次
- A Robust Unlearning Method with Adaptive Knowledge Guidance and Memory PreservationJingyuan Tian, Xiaofei ZhouAAAI 2026
