R-TOFU: Unlearning in Large Reasoning Models
Sangyeon Yoon, Wonje Jeung, Albert No
摘要
Large Reasoning Models (LRMs) embed private or copyrighted information not only in their final answers but also throughout multistep chain-of-thought (CoT) traces, making reliable unlearning far more demanding than in standard LLMs. We introduce Reasoning-TOFU (R-TOFU), the first benchmark tailored to this setting. R-TOFU augments existing unlearning tasks with realistic CoT annotations and provides step-wise metrics that expose residual knowledge invisible to answer-level checks. Using R-TOFU, we carry out a comprehensive comparison of gradient-based and preference-optimization baselines and show that conventional answer-only objectives leave substantial forget traces in reasoning. We further propose Reasoned IDK, a preferenceoptimization variant that preserves coherent yet inconclusive reasoning, achieving a stronger balance between forgetting efficacy and model utility than earlier refusal styles. Finally, we identify a failure mode: decoding variants such as ZeroThink and LessThink can still reveal forgotten content despite seemingly successful unlearning, emphasizing the need to evaluate models under diverse decoding settings. Together, the benchmark, analysis, and new baseline establish a systematic foundation for studying and improving unlearning in LRMs while preserving their reasoning capabilities. We release R-TOFU and code at https://ai-isl.github.io/r-tofu . LessThink ZeroThink Question Question
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning FailuresSangyeon Yoon, Hyesoo Hong, Wonje Jeung, Albert NoICLR 2026 · 被引用 3 次
- CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference OptimizationJunyi Li, Yongqiang Chen, Ningning DingACL 2026 · 被引用 1 次
- SEPS: A Separability Measure for Robust Unlearning in LLMsWonje Jeung, Sangyeon Yoon, Albert NoEMNLP 2025
它引用的顶会 Paper13
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 被引用 365 次
- In-Context Unlearning: Language Models as Few-Shot UnlearnersMartin Pawelczyk, Seth Neel, Himabindu LakkarajuICML 2024 · 被引用 217 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsZhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian 等ICLR 2024 · 被引用 98 次
- Unlearn What You Want to Forget: Efficient Unlearning for LLMsJiaao Chen, Diyi YangEMNLP 2023 · 被引用 40 次
相关 Paper
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language ModelsHongji Li, Manjiang Yu, Junchi Yao, PRIYANKA SINGH 等CVPR 2026 · 被引用 3 次
- STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning ModelsJingjing Zhou, Gaoxiang Cong, Li Su, Liang LiAAAI 2026
- Leak@: Unlearning Does Not Make LLMs Forget Under Probabilistic DecodingHadi Reisizadeh, Jiajun Ruan, Yiwei Chen, Soumyadeep Pal 等ICML 2026
- Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning SkillsChangsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia 等EMNLP 2025
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning StepsMartin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan BelinkovEMNLP 2025 · 被引用 37 次
